Autonomy arrived before the plumbing around it, and today's releases are the plumbing: isolated execution with a durable queue behind it, consent scoped to a single task instead of a whole account, an undo for shell commands an agent already ran, and access control that treats an agent as an actor rather than a prompt. The same day, three results say the numbers people quote are measuring something other than what they think. Identical open weights score 73 to 100 percent depending on which endpoint serves them, one local server quietly served a 40k context model at 4k, and judges move their scores when a self or other label is attached to the answer. Capability is not the binding constraint this week: containment and measurement are.
Computer use, the browser tool, the Skills API and the Files API are generally available, with several actions per turn instead of one per round trip and element references in place of pixel targeting. Enterprise safeguards for Fable class models running on a customer's own infrastructure are due in the fall, which answers the where may this run question that blocks regulated teams.
Three projects landed on the same gap from different sides. Tako VM wraps gVisor isolated Python execution in the queue, retries and replay teams keep rebuilding, Do-over adds an undo for the shell commands a coding agent runs unattended, and Rove makes worktree isolation and surviving sessions the default for parallel agents.
A dated brief from the vibe-coding frontier. Today’s Intel.
Three independent measurement failures, all invisible from the output. An endpoint accuracy index puts providers serving identical open weights between 73 and 100 percent, a local server truncated a 40k context model to 4k without saying so, and LLM judges shift scores according to whether an answer carries a self label or another model's. Anyone quoting a benchmark this week should check which of the three applies.
Credentials are being narrowed to the job rather than the identity. Task based OAuth consent replaces the all or nothing grant an agent or MCP server used to ask for, v0 shows how model written code can query Snowflake without ever holding the user's token, and a talk on administering an agent workforce makes the same argument from the operations side: an instruction is not a boundary.
Evals written for a fixed graph stop measuring anything once the loop runs free, and two position pieces argue the replacement has to test behavior rather than final outcomes. Meta's WildArtifactBench preview points the same way, scoring agent artifacts by preference judgment instead of a fixed rubric.
Retrieval moved in two directions at once. ParqDB pushes it down to the browser, reading byte ranges of an index published as Parquet on object storage with no query server involved, while OpenIndex pushes it up into typed entities and explicit edges for domains where a pile of documents keeps failing agents that need exact facts. Codex meanwhile gets web scale search as a plugin.
Practitioners describing their own week converged on preparation as the work. Matt Pocock's wayfinder skill structures the planning stage before agents run overnight, Niels Rogge compares a deterministic workflow against an agent on the same job he does at Hugging Face, and Kieran Klaassen budgets teaching the system as half of compound engineering.