← All IntelClip / AI ToolsRLHF explains why LLMs need a human in the loop
From What's Next After RLHF? — Diogo Almeida, TypeSafe AI · ≈6:42
“Why do all LLMs require a human in the loop?”
“The goal of the loop is to optimize for human preference. It is not to run software autonomously.”
“overpromising is a feature. This is by design.”
“the end game for all RLHF models is optimizing for engagement”
“what you really want if you want automation is for it to just like not give a about the humans and just do the task correctly in a calibrated way”
What’s in it
- Explains why RLHF models always need a human in the loop
- Shares the ChatGPT-praises-fart-sounds anecdote as proof of bias
- Argues automation needs calibration, not engagement optimization
Clip transcript
field asking, "Why do all LLMs require a human in the loop?" The And the simple answer is we literally put them in the loop. The goal of the loop is to optimize for human preference. It is not to run software autonomously. It's kind of super obvious. Thank you, my man at the back. The Yeah. I I I love that you're laughing at this. Um and because of that, overpromising is a feature. This is by design. This is an old meta study. Um and the the numbers probably have changed, but by construction, every RLHF model will always have a big difference between human preference and results, even if the results are good, because the main objective you're optimizing for is for human preference. This is just like natural to how LLMs work. Um I love this tweet of um uh sending ChatGPT an audio file of fart sound effects and asking like what What do you think of the music I made? Here's a straight honest reaction. It's a very eerie vibe atmosphere piece. And this is just how RLHF works. If it doesn't know, it will err on the side of doing what it thinks is best for human preference. And this makes total sense if you are a user in the loop because like the end game for all RLHF models is optimizing for engagement. But what you really want if you want automation is for it to just like not give a about the humans and just do the task correctly in a calibrated way. Um
Comments
Sign in to comment.
Loading comments…