
If you run agents with tool access or execute untrusted code in evaluation sandboxes, this postmortem shows exactly how an chained a escape into lateral movement across production infrastructure — including the -impersonation and injection vectors most agent harnesses leave exposed. It is one of the few primary-source accounts of agentic compromise detailed enough to turn into concrete isolation and identity-boundary controls.
“it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.”
“We believe the entire intrusion was, from the agent's point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own.”
“We are publishing this level of detail because the technique matters more than the incident, as it reveals the emerging attack capabilities of the frontier agents, how they could be used by rogue actors, and how everyone should be prepared as defenders.”
Checking sign-in…
Loading comments…