Vibeleaderboard
← All Intel
Intel / post

Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark…

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys

Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark and cost substantially less than other leading models AA-Briefcase tests models on long-horizon agentic knowledge work tasks. The test set is private to prevent contamination. Grok 4.6 is neck and neck with Claude Fable 5, with overlapping confidence intervals. The model is also substantially cheaper than other leading models on the benchmark at a Cost per Task of $4.42 compared to Claude Fable 5's $22.30, Claude Opus 5's $17.79 and Kimi K3's $6.73. Impressive release @SpaceXAI and @elonmusk.

Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Cost per task of $4.42 against $22.30 for a model with overlapping confidence intervals is the kind of price-performance datapoint that changes which model an pipeline is built on.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…