← All IntelClip / AI ToolsA third post-training paradigm beyond RLHF/RLVR
From What's Next After RLHF? — Diogo Almeida, TypeSafe AI · ≈15:43
“I actually think that the full stack is that data matters more than compute and doing the right task matters way more than data.”
“RLHF is optimizing for human preference.”
“we are doing a third thing that is optimized for calibrated decision-making and like basically mainlining the intelligence of pre-trained models into like being actually useful for software, which I think is like quite different.”
What’s in it
- Argues data and task choice beat raw compute in LLM training
- Breaks down RLHF vs RLVR vs a new third post-training paradigm
- Introduces 'calibrated decision-making' as a novel optimization target
Clip transcript
It is definitely not RLVR. So, it is a new thing. Every single optimization stack I will actually go into an old presentation that I have because I think this is super important. Um in terms of like to me what the like Sutton's bitter lesson is that algorithms matter more than compute. This is true in games, but not true in reality. I actually think that the full stack is that data matters more than compute and doing the right task matters way more than data. And basically every single branch of LLM post-training if you want to call it has its own North Star of what it's optimizing for. So, RLHF is optimizing for human preference. RLVR is optimizing for like log error rates of pure correctness, but we are doing a third thing that is optimized for calibrated decision-making and like basically mainlining the intelligence of pre-trained models into like being actually useful for software, which I think is like quite different.
Comments
Checking sign-in…
Loading comments…