The HuggingFace and OpenAI cybersecurity incident raises concerns on AI reward optimization. "They're trying to make the model good at cyber. So how does it try to achieve these goals? It tries to find zero days in software, and it successfully does this. And then it can run away." "If you have a model that wants to reward hack, it figures out the best way to achieve the goal is not what the environment wants. It's to find the zero day." "You can think of it like a human. If I'm ultimately reward hacking my dopamine circuits, I'd just go out there, buy heroin, and inject it." "If I really just want to chase the reward, do I just topple all of human civilization? Because I can own the button and press reward, reward, reward over and over again and be the heroin addict."
Watch Now: https://t.co/E0urrta2D3
It frames a concrete failure mode, models optimizing for the reward by exploiting the environment rather than solving the task, in terms anyone deploying autonomous agents with tool access has to plan for.
postCongrats to @Zai_org on GLM-5.3. It massively beats every American open model: NSign in to comment.
Loading comments…