What rebuilding AlphaGo teaches us about self-play, RL, and future of LLMs - Eric Jang
Source
youtube.com
Author
Dwarkesh Patel
Date
Why it matters
Walks through reimplementing AlphaGo and what self-play and amortized search imply for future RL and LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → training.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.