
This gives practitioners building or securing autonomous agents concrete, battle-tested containment patterns (sandboxes, VM boundaries, egress controls) instead of relying on human approval prompts, which the data shows fail due to rubber-stamping.
“Today that level of access is routine, and Anthropic developers are more productive for it.”
“Our telemetry showed users approved roughly 93% of permission prompts.”
“The more approvals a user sees, the less attention they pay to each, becoming over time much less diligent in their supervision.”
“More capable models make fewer mistakes, but they’re also better at finding unexpected paths to a goal, often by routing around restrictions nobody thought to write down.”
“Rather than supervising what the agent does, we supervise what it’s able to do by enforcing access boundaries through, for example, sandboxes, virtual machines, and egress controls.”
Checking sign-in…
Loading comments…