Cloudflare Containers, rebuilt to scale agent sandboxes
Source
blog.cloudflare.com
Date
Key takeaways · AI-distilled
Because each sandboxAn isolated environment where AI-generated code or agent actions run without being able to touch anything real.Full definition → picks its image in code at start time, rollouts need no platform config: Cloudflare suggests canarying a toolchain on 5% of new sandboxes by hashing the Durable Object ID, pinning active projects to their image, and rolling back by changing future starts.
On ComputeSDK's Burst TTI benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → of 100 concurrent sandboxes, median startup fell from 4.049s to 648ms and p99 from 6.717s to 1.129s. Cloudflare credits placement near the Durable Object, hosts with the image cached, and restoring prepared VMs instead of booting fresh.
Snapshots are immutable and reusable, so one snapshot can seed many isolated sandboxes from the same baseline. Cloudflare pitches this for evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → comparing prompts, models or AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → versions without environment drift, and for RL coordinators that fork and grade attempts.
Cloudflare's migration guidance warns that a snapshot restores only on the image it was taken from, so an image deploy can lock users out of workspaces whose files live only in snapshots. It says to keep file-level backups as the source of truth.
New features are available only through ctx.container. The Container class and legacy Sandbox class are maintained through December 31, 2026 and then stop getting updates; migrating usually means changing extends Container to extends DurableObject.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
sandbox — An isolated environment where AI-generated code or agent actions run without being able to touch anything real.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Agents can create sandboxes per task that start in under a second and pause and resume, which matters for anyone running code-executing agents at scale.