← All IntelClip / AI ToolsHallucination as intrinsic to preference optimization
From What's Next After RLHF? — Diogo Almeida, TypeSafe AI · ≈14:19
“I actually don't think that pre-training is the problem.”
“I think pre-training is uh phenomenal.”
“There's an asymmetry in the reward model like what GANs have that allow for um that encourage the models to drop modes and be confident because it's very easy to see when the model is not confident and to punish that from a reward model perspective.”
What’s in it
- Argues pre-training isn't the real cause of AI hallucination
- Frames hallucination as a GAN-like mode-collapse problem in reward models
- Explains why reward models punish honest uncertainty and reward false confidence
Clip transcript
suggesting. Um I will say that that's complicated. And I actually think I don't have the time to answer that particular question. I will give like my simplified view on this. And it the answer is I actually don't think that pre-training is the problem. I think pre-training is uh phenomenal. Like the fact that we compress the knowledge of the internet into like this core of intelligence that then can be utilized is incredible. And the pre-trained models are incredibly intelligent. Uh and I believe that the problem is like how we unearth it. And hallucination [clears throat] to me is intrinsic to um optimizing for human preference. Like there's an asymmetry in the reward model kind of like a GANs have. Oh, I really should not get This is a very advanced topic. But there's an asymmetry in the reward model like what GANs have that allow for um that encourage the models to drop modes and be confident because it's very easy to see when the model is not confident and to punish that from a reward model perspective. It's very complicated, but
Recommended reads
articleUnified Hallucination Fuzzing for Multimodal Large Language ModelsPengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You
articleExtrinsic Hallucinations in LLMsLilian Weng
articleOpen challenges in LLM researchChip Huyen
Comments
Sign in to comment.
Loading comments…