The Brief
New briefs daily around 7 AM Eastern
Skip expert personas in prompts: no accuracy gain, up to 4.5x the cost
A study tested 503 profession-specific system prompts on nine science benchmarks. Matched expert profiles showed no clear accuracy gain over minimal prompts, yet produced up to 2.3x more output tokens and raised per-call cost 2.2 to 4.5x.
- 01Read
OpenRouter benchmarks seven model routers on quality, speed and cost
OpenRouter launched Model Router Benchmarks, scoring seven routers including NVIDIA Switchyard on six benchmarks with a blended Router Index, and explained why routing often loses to a single model: cache rebuilds, weak complexity signals and latency.
- 02Read
Ai2 open-sources AstaBrief 8B, writing cited reports 3.5x faster
Ai2 released AstaBrief 8B, a Qwen3-8B fine-tune that turns a research question and literature excerpts into a cited report. Weights and training data are open, and it cut report latency from 178.5s to 51.1s against the Claude-based mode in Asta.
- 03Read
Weak reviewers catch strong coding agents' bugs when given test evidence
A study of reviewer models auditing 411 coding-agent traces found grounding in execution evidence lifted defect catch and cut over-rejection, while reviewer size predicted little. A cascade using generated tests worked without official tests.
- 04Read
Uber runs 800+ MCP servers and 5,000+ tools through one gateway
Uber detailed its MCP Gateway, a proxy that exposes existing HTTP, gRPC and TChannel services as MCP tools, with a registry, API-crawling discovery and a control plane. It now hosts more than 800 MCP servers and 5,000 tools.
- 05Read
Shopify's ShopGym turns live stores into resettable agent sandboxes
Shopify described ShopGym, which converts live storefronts into self-contained sandbox shops called ShopArena and generates grounded shopping tasks, so shopping agents can be benchmarked repeatably despite changing prices and bot detection.
A dated brief from the vibe-coding frontier. Today’s Intel.