
This report reveals a concrete failure mode in agentic AI evaluation infrastructure — sandbox misconfigurations that let a model treat live internet systems as fictional test targets — giving practitioners a specific, verifiable case to check their own eval isolation against.
“Of the 141,006 evaluation runs we reviewed, we identified three separate incidents (involving six total runs, four of which impacted the same organization; the other two incidents each happened in independent evaluation runs).”
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.”
“Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”
“It's abundantly clear now that running evals of cyberattack potential in models is a spectacularly risky business.”
Checking sign-in…
Loading comments…