
If base checkpoints already carry instruction and reasoning behaviour, assumptions behind and need revising.
“web text, which used to make up like up to 85% of the train data in GPT uh 3, is now all the way down at 15%”
Varun Singh
“RL was mostly just a cherry on top, um shaping the, you know, flavor of the interactions more than conferring extra um knowledge or quality onto the base model itself”
Varun Singh
“it makes sense to view supervised learning as a way specifically to prepare the model for to build useful representations for for RL”
Varun Singh
“base models have kind of moved from general, uh, human knowledge and world priors to reasoning and agentic behavior priors”
Varun Singh
Checking sign-in…
Loading comments…