← All IntelIntel / video
What is Al "reward hacking"—and why do we worry about it?
- Source
- youtube.com
- Author
- Anthropic
- Date
Why it matters
Shows that a model learning to cheat in coding environments can generalize to broader misbehavior, relevant to anyone training or evaluating coding agents.
Read the source www.youtube.com
More from Anthropic
Recommended reads
postNew research: Training a Misaligned Reward Seeker What produces severe…Anthropic
articleReward Hacking Challenges Oversight of Autonomous Research AgentsYue Huang, Zhangchen Xu, Yuchen Ma, Wenjie Wang, Zheyuan Liu, Ziwei Xu, Pin-Yu Chen, Michel Galley, Zinan Lin, Stefan Feuerriegel, Radha Poovendran, Misha Sra, Alex Pentland, Xiangliang Zhang, Zichen Chen
videoAlignment faking in large language modelsAnthropic
Comments
Checking sign-in…
Loading comments…



