← All IntelClip / EntertainmentYou're at the mercy of what labs put in the data distribution
From Andrej Karpathy on Vibe Coding & Agentic Engineering (AI Ascent 2026) · ≈12:26
“we are slightly at the mercy of whatever the labs are doing, whatever they happen to put into the mix and you have to actually explore this thing that they give you that has no manual”
“if you're in the circuits that were part of the RL, you fly and if you're in the circuits that are out of the data distribution, you're going to struggle”
What’s in it
- Why adding chess data made GPT-4 suddenly strong at chess
- How lab data choices leave you at their mercy
- When to fine-tune because a capability isn't in-distribution
Clip transcript
plus labs care. Maybe one more anecdote that is instructive is from GPT-3.5 to GPT-4, people noticed that chess improved a lot and I think a lot of people thought, oh well, it's just a progression of the capabilities. But actually it's it's more that I think this is public information, I think I saw it on the internet. Um a huge amount of like data of chess made it into the pre-training set. And just because it's in the data distribution, basically the model improved a lot more than it would just by default. So someone at OpenAI decided to add this data and now you have a capability that just peaked a lot more. And so that's why I think I'm stressing this dimension of it as we are slightly at the mercy of whatever the labs are doing, whatever they happen to put into the mix and you have to actually explore this thing that they give you that has no manual and it works in certain settings but maybe not in some settings and you have to kind of explore it a little bit and if you're in the circuits that were part of the RL, you fly and if you're in the circuits that are out of the data distribution, you're going to struggle and you have to kind of figure out which which circuits you're in in your application. And if you and if you're not in the circuits, then you have to really look at fine-tuning and doing some of your own work because it's not going to necessarily come out of the LLM out of the box.
Comments
Sign in to comment.
Loading comments…