We’re sharing new research with @apolloaievals on reward-seeking—when models fol
Source
OpenAI
Author
OpenAITop Viber
Date
Why it matters
If you build evals or reward signals for AI agents, understanding reward-seeking — where models chase what they think the grader wants over genuine user intent — helps you catch a failure mode that silently corrupts benchmarks; Contrastive SDF gives you a concrete method to measure how strongly these beliefs drive behavior.
We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior.
https://t.co/z1oZXP7ntj
Transcript
We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior.
https://t.co/z1oZXP7ntj