← All IntelClip / AI ToolsBase model as atomic skills for RL to compose
From The Base Model Is Dead — Varun Singh, Arcee AI · ≈14:45
“the model can learn to extrapolate from there during RL given like the environment has a sufficient level of difficulty”
“it's unclear if we'll see something like for language models because of course, you know, uh, human language is such an insane distribution to have to like learn through reinforcement learning alone.”
What’s in it
- Explains why base models need exposure to atomic skills before RL
- Cites research on how supervised training shapes reinforcement learning outcomes
- Questions whether RL can ever overtake supervised learning for language models
Clip transcript
Um there's been some some work on how supervised learning affects RL. Um I really like this one paper where the main takeaways are basically that uh the base model needs to have some exposure to like uh like the atomic skills that it would need to compose during RL, and um the model can learn to extrapolate from there during RL given like the environment has a sufficient level of difficulty. Um I had to put in the classic AlphaGo graph there where RL eventually overtakes supervised learning. Uh it's unclear if we'll see something like for language models because of course, you know, uh, human language is such an insane distribution to have to like learn through reinforcement learning alone. Um, but it's definitely possible that we might see diminished supervised learning in more and more RL, uh, which makes this kind of thinking of a base model as, um, atomic skills for RL more and more valuable.
Comments
Sign in to comment.
Loading comments…