Vibeleaderboard
← All Intel
Intel / post

Failure analysis: We classified 1,567 failing AA-AnalystAgent attempts across…

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys

Failure analysis: We classified 1,567 failing AA-AnalystAgent attempts across ten leading models into seven failure modes, tagging each attempt with every mode it exhibited. The most widespread is anchoring on a wrong early hypothesis, present in 57% of failures. Models commit to a source or interpretation early and defend it for the rest of the trajectory. The sharpest contrasts are in how models treat sources. Gemini 3.1 Pro Preview takes sources at their word but struggles with execution. It sits below the ten-model median on both misreading domain terms and overriding evidence in favor of a generic prior, while posting above-median rates of modeling, scaling and aggregation errors (51% of its failures) and skipped verification (39%). @SpaceXAI's Grok 4.5 (high) is the reverse. It understands the field's language, then substitutes its own assumptions for what the documents say, misreading domain terms on 23% of its failures against a 38% median, while overriding evidence on 54% against a 43% median. Kimi K3 (max) misreads on 48%, 25 points clear of Grok 4.5.

Terms in this piece · Glossary
  • guardrails — The checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.
Why it matters

Knowing that most failures come from early anchoring, not arithmetic, tells you where to spend : force source re-reading and hypothesis revision rather than more verification passes. The per-model split also explains why swapping models changes which failures you see.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…