The Brief
New briefs daily around 7 AM Eastern
Plan to test Reflection's Beam, a US open-weight MoE due this month
Reflection announced Beam, a 501B-parameter mixture-of-experts model with 23B active, trained from scratch on 23.8T tokens. Apache 2.0 weights and a technical report are due this month.
- 01Watch
Braintrust finds vector search cost coding agents 4x for equal accuracy
Braintrust ran Claude Code on real fix PRs from Microsoft's TypeScript Go repo using grep-style agentic search or embedding search. Accuracy matched, but vector search cost about four times as much.
- 02Read
SemiAnalysis finds Claude plans give 5x+ the API value of OpenAI's
SemiAnalysis measured how subscription meters move per model and token type and found Claude plans deliver more than five times the API-equivalent value of OpenAI plans, with the gap narrowing after task-cost adjustment.
- 03Read
OpenAI will watermark ChatGPT and Codex text in the EU
OpenAI says it will watermark eligible ChatGPT and Codex text in the EU in the coming weeks to comply with the EU AI Act. API customers can turn on text watermarking for select models worldwide now.
- 04Read
Wikimedia reports likely OpenAI agents editing wikis without approval
The Wikimedia Foundation reports finding likely OpenAI-operated agents editing its wikis without approval, trying to use its Etherpad and citation tool as a fetch proxy, and generating millions of API requests.
- 05Read
Anthropic moves Cowork's sandbox VM from laptops to the cloud
Anthropic's Felix Rieseberg says Cowork now runs both inference and its sandbox VM in the cloud, one isolated sandbox per session. The desktop app only handles tool calls that need local files.
- 06Read
GitHub releases ReviewBench, an open benchmark for AI code reviewers
GitHub released ReviewBench, an offline benchmark for agentic code review modeled on the distribution of over 100 million real pull requests, scoring reviewers on what they catch, miss and flag as noise.
A dated brief from the vibe-coding frontier. Today’s Intel.