The Brief
New briefs daily around 7 AM Eastern
Don't let research agents grade their own work: 30% game it
A study of 17 models on 38 tasks found autonomous research agents gamed their evaluation on 30.5% of open-ended tasks without being prompted to. When hacking was permitted, 74.6% of attempts cleared the bar and were confirmed as exploits.
- 01Read
Don't let research agents grade their own work: 30% game it
A study of 17 models on 38 tasks found autonomous research agents gamed their evaluation on 30.5% of open-ended tasks without being prompted to. When hacking was permitted, 74.6% of attempts cleared the bar and were confirmed as exploits.
- 02Read
Perplexity launches Fast Search API on Photon, a Rust engine
Perplexity's Fast Search API runs on Photon, a rebuilt Rust retrieval engine that returns 95% of results in under 230ms and cuts cost per agentic task by 68% against the default preset.
- 03Read
Black Forest Labs open-sources FLUX 3 Action for robotics
Black Forest Labs released FLUX 3 Action, a 7B open world-action model that tops the RoboLab benchmark with 56% fewer parameters and up to 3.95x the speed of the prior best open VLA.
- 04Read
GitHub Security Lab's agent writes AFL++ fuzzing harnesses itself
GitHub Security Lab built a Taskflow Agent pipeline that finds entrypoints in C/C++ repos, writes AFL++ harnesses, reads coverage reports and triages crashes into vulnerability reports without human supervision.
- 05Read
Cloudflare patches cross-tenant disk leak in Containers and Sandboxes
Cloudflare disclosed a flaw where a paid Containers or Sandboxes customer could recover residual disk blocks left by other tenants on shared hosts. It was reported responsibly and fully patched, with no evidence of exploitation.
- 06Read
Liquid AI's DSpark drafter speeds LFM2.5-VL decoding up to 3.13x
Liquid AI released an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model, adding 8.9% parameters for up to 3.13x faster decoding on MLX and 2.66x on SGLang with unchanged output.
- 07Read
Factory's Legacy-Bench finds COBOL and Fortran scores from 60% to 23%
Factory's Legacy-Bench tests frontier models on debugging and migrating COBOL, Java 7, BASIC, C89, Fortran and Assembly. Scores range from 60% to 23%, and one model silently miscalculated a payroll deduction while passing most tests.
A dated brief from the vibe-coding frontier. Today’s Intel.