Vibeleaderboard
Archive — Daily Brief← Intel

The Brief

The Brief · Wed, Aug 5Auto-synthesized · Cited · 34 sources

New briefs daily around 7 AM Eastern

Evals are the failure mode: today's research says the score isn't measuring the risk

The day's research converges on a single uncomfortable finding: the numbers we use to approve, rank, and ship models are frequently measuring something other than what we think. Cost-sensitive policy text embedded in a code-review prompt shifts reported failure probabilities by 13-17 points on identical evidence; clinician preference votes rank unsafe clinical outputs highly; up to 42% of correct legal answers cite the wrong statute or none at all; and permission-aware access control in on-device memory assistants fails across every system tested. The shared repair is decomposition — separate risk elicitation from policy, score answer and authority jointly, diagnose root cause instead of retrying — and the same instinct shows up in engineering, where a single structured intermediate representation beat multi-agent repair loops at 8-39x fewer LLM calls. Meanwhile the platform layer kept building for agents that nobody can yet audit, which is why the security threads of the day are worth reading alongside the papers, not after them.

Generated 09:16 UTC · from the corpus, not a 30-day windowAsk the brain →

A dated brief from the vibe-coding frontier. Today’s Intel.