What's Next After RLHF? — Diogo Almeida, TypeSafe AI
www.youtube.com- Category
- AI Tools
- Type
- ARTICLE
- Added
- Aug 4, 2026
About
RLHF made models that are extraordinary at pleasing the human in the loop, and Diogo Almeida, a GPT-4 co author, argues that is exactly the problem. Optimizing for human preference optimizes for engagement and for overpromising, the same pressure that makes a model confidently agree that a fart audio file is a symphony. That produces two camps: one where models act as assistants with a human catching mistakes, where RLHF shines, and one where they operate autonomously with real stakes, where the
Why it made the leaderboard
Sycophancy and overclaiming are not bugs on top of RLHF; the argument is that they are what it optimizes for.
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.