
Karpathy builds a GPT from an empty file to training on real text, narrating every decision. Still the clearest way to replace "LLMs are magic" with an actual mental model of what is happening under the hood.
“While the GPT-2 (124M) model probably trained for quite some time back in the day (2019, ~5 years ago), today, reproducing it is a matter of ~1hr and ~$10.”
“We basically start from an empty file and work our way to a reproduction of the GPT-2 (124M) model.”
“The git commits were specifically kept step by step and clean so that one can easily walk through the git commit history to see it built slowly.”
Checking sign-in…
Loading comments…