
It measures reliability across repeated attempts rather than a single lucky pass, which is the property that decides whether an agent removes work or adds verification. The is open source, so the setup is reproducible.
“In real analyst roles, professional judgment and expertise are as important raw quantitative capabilities. AA-AnalystAgent tests this and requires models to interpret sources, decide which exceptions and caveats apply, and settle on a methodology to successfully complete a task.”
ArtificialAnlys
“Because handing analyst tasks to agents requires not just correct answers but consistent ones, AA-AnalystAgent runs each task five times and reports pass^5 as its headline metric (we call this ‘pass-all-5’). pass^5 means models must get a task correct every time it tries across five independent attempts to pass.”
ArtificialAnlys
“Reliability separates the top of the leaderboard more than raw capability: GPT-5.5 (xhigh) has the highest pass@1, but Opus 5 leads on pass^5 because it repeats what it gets right.”
ArtificialAnlys
“Committing early to a wrong interpretation is the most widespread way models fail, appearing in 57% of the failures we classified.”
ArtificialAnlys
“The price of a given score varies enormously: Claude Sonnet 4.6 and @Xiaomi's MiMo-V2.5-Pro both score 20%, at $1.34 and $0.05 per task.”
ArtificialAnlys
postDeepSeek V4 Pro 0813 scores 53 on the Artificial Analysis Intelligence Index, 8
postGoogle has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash aSign in to comment.
Loading comments…