Vibeleaderboard
Archive — Daily Brief← Intel

The Brief

The Brief · Fri, Aug 7Auto-synthesized · Cited · 27 sources

New briefs daily around 7 AM Eastern

Benchmark headlines break down: cost per task, not per token, decides the day

The day's most useful number was not a leaderboard rank but a denominator. Artificial Analysis showed Qwen3.8 Max cutting per-token price while more than doubling cost per Intelligence Index task, winning GDPval-AA Elo largely by taking roughly four times as many turns, and regressing ten points on AA-Omniscience as it stopped abstaining — three different ways a headline score can hide latency, spend and hallucination risk, arriving alongside an Intelligence Index grader change that makes cross-version comparisons invalid. Against that, Meta's Muse Spark 1.2 landed on the cost-per-task frontier and Kimi K3 reached GA in Copilot before a rollout pause, giving builders cheaper defaults whose real economics still need per-workload measurement. Underneath the model news, the plumbing converged: Agent Plugins shipped as a cross-client standard from OpenAI and Cursor on the same day the MCP ecosystem moved toward stateless servers and web-side tool interfaces.

Generated 03:14 UTC · from the corpus, not a 30-day windowAsk the brain →

A dated brief from the vibe-coding frontier. Today’s Intel.