← All IntelClip / AI ToolsMid-training as a distinct phase
From The Base Model Is Dead — Varun Singh, Arcee AI · ≈12:44
“Um a lot of models though are training with much longer context in pre-training and there's no reason that these data sets can't be pulled back into the mix to allow for more stable representations from the very beginning.”
What’s in it
- Explains what 'mid-training' actually does to prep base models for RL
- Shows why agentic traces are getting mixed into long-context training runs
- Argues pre-training data can be reused later for more stable models
Clip transcript
Another uh interesting thing that um is changing in base models now is that you is this whole advent of mid-training, uh which is exposing the model to the distribution that it would see during post-training in RL and at a longer context, so for things like agentic traces to be allowed into the mix and uh to kind of help prepare the model that way. Um a lot of models though are training with much longer context in pre-training and there's no reason that these data sets can't be pulled back into the mix to allow for more stable representations from the very beginning.
Comments
Checking sign-in…
Loading comments…