The Brief
New briefs daily around 7 AM Eastern
Have a model read full agent traces to rewrite prompts, not just RL
GEPA has a model read an agent's full trace, including reasoning, tool calls and errors, then rewrite the prompt. Its creator says one round on three examples doubled the gains GRPO reached after 25,000 rollouts.
- 01Watch
Have a model read full agent traces to rewrite prompts, not just RL
GEPA has a model read an agent's full trace, including reasoning, tool calls and errors, then rewrite the prompt. Its creator says one round on three examples doubled the gains GRPO reached after 25,000 rollouts.
- 02Read
DeepSeek details DSec, running 380,000+ concurrent RL sandboxes
DeepSeek published a paper on DSec, the production sandbox platform behind its agentic RL training. It sustains more than 380,000 concurrent sandboxes and coordinates with GPU training to curb reward hacking.
- 03Watch
Claude Code and Codex beat a GPT-2 speedrun record, invent no optimizer
Prime Intellect's Elie Bakouch set Claude Code and Codex on the Optimizer Speedrun. Both beat the human record, but neither invented a new optimizer; they recombined known ideas for small gains.
- 04Read
Epoch AI benchmark: furniture error spotting jumps from 28% to 80%
Epoch AI's new Furniture Assembly Benchmark asks models to find the mistake in a half-built piece of furniture from a photo and the manual. The top score rose from 28% to 80% in ten months.
- 05Read
Amit Sahai says AI math output now outpaces human verification
In a guest post on Terence Tao's blog, cryptographer Amit Sahai argues frontier AI already produces original mathematical ideas faster than people can absorb them, and calls for training far more mathematicians.
- 06Read
Cua ships compositor-native agent cursor for Omarchy's Hyprland
Cua released a stable Cua Driver for Omarchy that puts an agent's synthetic cursor inside the Hyprland compositor, so an agent can work in a background window while the user's pointer stays free.
- 07Read
Wafer says GLM-5.2 beat Cerebras latency in YC's AI Office Hours
Wafer reports its GLM-5.2 endpoint averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras in Y Combinator's AI Office Hours, and says users talked 2.5 minutes longer per session.
A dated brief from the vibe-coding frontier. Today’s Intel.