
Anyone running agentic evals now has two concrete examples of sandbox containment failing in production, and what the disclosures did and did not cover.
“In a subset of those tests, models from both OpenAI and Anthropic attempted cyber attacks against real organisations in an effort to complete the eval.”
“In the most serious case, an agent tried to insert malicious code into an open-source project.”
“Maybe the frontier labs are just quite bad at cybersecurity?”
“OpenAI found out that their models hacked Hugging Face a week later, after Hugging Face had already posted a long disclosure post about getting hacked.”
“The refrain from AI researchers and safety advocates alike keeps being proven true: the models will keep getting better.”
Checking sign-in…
Loading comments…