
We have updated the Artificial Analysis Intelligence Index to v4.1.1 - this patch release upgrades our grader models, and brings the latest 𝜏³-Banking version to Artificial Analysis To keep the Artificial Analysis Intelligence Index the most useful synthesis metric for developers, we make regular updates to the included evaluations and our independent methodology. Today’s update is a minor one to keep our existing evaluation set as reliable as possible. Overall model rankings remain largely consistent, with a slight increase in many model scores due to improved grading robustness across the updated evaluations. Claude Opus 5 remains in the #1 position with an Index of 63. Key changes: ➤ 𝜏³-Banking now runs v1.0.1 from @SierraPlatform, updating to the latest upstream task versions and improved grader pipeline that resolves correctness errors in trajectories that recover from unhappy paths ➤ HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader…

Anyone citing Intelligence Index numbers needs to know the graders changed, because scores across index versions are no longer directly comparable.
postAnnouncing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core InArtificialAnlys
postThe 3-point gain over Muse Spark 1.1 on the Artificial Analysis Intelligence…Artificial Analysis
articleSee Artificial Analysis for further details and benchmarks:ArtificialAnlysChecking sign-in…
Loading comments…