← All IntelClip / AI ToolsWhy models always 'look right' regardless of correctness
From What's Next After RLHF? — Diogo Almeida, TypeSafe AI · ≈8:11
“Um Lesson number two is that today's AI was designed for assistance through optimizing for human preference.”
“no matter how wrong the models are, they will look right because of the asymmetry within the reward model in RLHF.”
“this is where a lot of like the dilemma in the field stems from because people really want automation to happen.”
What’s in it
- Explains why RLHF-trained AI models look confident even when wrong
- Unpacks the human-preference optimization flaw baked into today's AI
- Frames the core tension driving debates over AI automation
Clip transcript
calibrated way. Um Lesson number two is that today's AI was designed for assistance through optimizing for human preference. This is like it's like in the name. This is not like a controversial take. And the consequences are maybe more controversial, but it's like very obvious if you think about what we really are optimizing for, which is no matter how wrong the models are, they will look right because of the asymmetry within the reward model in RLHF. Um and this is where a lot of like the dilemma in the field stems from because people really want automation to happen.
Comments
Checking sign-in…
Loading comments…