Vibeleaderboard
← All Intel
Intel / video

What rebuilding AlphaGo teaches us about self-play, RL, and future of LLMs - Eric Jang

Source
youtube.com
Author
Dwarkesh Patel
Date
Why it matters

Walks through reimplementing AlphaGo and what self-play and amortized search imply for future RL and training.

Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Read the source www.youtube.com
More from Dwarkesh Patel
Recommended reads
Comments

Checking sign-in…

Loading comments…