Vibeleaderboard
← All Intel
Intel / post

Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core In

Source
ArtificialAnlys
Date
ArtificialAnlys@ArtificialAnlys

Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations The Capability Indices map tasks from O*NET occupations to benchmarks that represent them, weighting each benchmark by how often its capability appears across the work. We cover six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics Key changes: ➤ Finance & Accounting: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Finance), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Strategy & Ops: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Legal: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context…

Read the full post on X

Context

Artificial Analysis builds 'Capability Indices' that try to score AI models on real job tasks rather than abstract puzzles: they take the task list behind specific U.S. occupations from the O*NET database, match each task to a that represents it, and weight benchmarks by how often that kind of work actually appears in the job.

In version 1.1, Artificial Analysis added a new agentic-tool-use benchmark to the Finance, Strategy and Ops, Legal, and Healthcare indices, dropped an agentic-customer-interaction benchmark from all of them, and added long- and document-based tests to several. Because the indices are reweighted rather than just extended, a model's rank in a given profession's index can shift even if its own scores haven't changed.

Under the new weighting, Artificial Analysis reports Claude Fable 5.1 still leads all six indices, with GPT-6 Astra second in Finance and Accounting, Strategy and Ops, Legal, and Engineering, while models take narrower wins: Kimi K3 in Finance and Accounting, Legal and Economics, DeepSeek V4.1 Flash in Strategy and Ops, and GLM-5.3 in Healthcare and Engineering.

Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
  • tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…