Failure analysis: We classified 1,567 failing AA-AnalystAgent attempts across…
- Source
- Artificial Analysis
- Date

Failure analysis: We classified 1,567 failing AA-AnalystAgent attempts across ten leading models into seven failure modes, tagging each attempt with every mode it exhibited. The most widespread is anchoring on a wrong early hypothesis, present in 57% of failures. Models commit to a source or interpretation early and defend it for the rest of the trajectory. The sharpest contrasts are in how models treat sources. Gemini 3.1 Pro Preview takes sources at their word but struggles with execution. It sits below the ten-model median on both misreading domain terms and overriding evidence in favor of a generic prior, while posting above-median rates of modeling, scaling and aggregation errors (51% of its failures) and skipped verification (39%). @SpaceXAI's Grok 4.5 (high) is the reverse. It understands the field's language, then substitutes its own assumptions for what the documents say, misreading domain terms on 23% of its failures against a 38% median, while overriding evidence on 54% against a 43% median. Kimi K3 (max) misreads on 48%, 25 points clear of Grok 4.5.


- guardrails — The checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.
Knowing that most failures come from early anchoring, not arithmetic, tells you where to spend : force source re-reading and hypothesis revision rather than more verification passes. The per-model split also explains why swapping models changes which failures you see.
postAnnouncing AA-AnalystAgent, our new agentic benchmark for quantitative analysis…Artificial Analysis
articleThe Knowing-Saying Gap: When Probes See Errors that Confidence MissesJyotin Goel, Ipshita Bandyopadhyay, Justin Shenk
articleSearchAuditor: Auditing and Attributing Failures in Long-Horizon Search AgentsZhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
Checking sign-in…
Loading comments…


