Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core In
- Source
- ArtificialAnlys
- Date

Announcing Artificial Analysis Capability Indices v1.1, updated with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations The Capability Indices map tasks from O*NET occupations to benchmarks that represent them, weighting each benchmark by how often its capability appears across the work. We cover six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics Key changes: ➤ Finance & Accounting: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Finance), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Strategy & Ops: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context Reasoning, and AA-Briefcase added to Agentic Knowledge Work ➤ Legal: Agentic Tool Use added as a new capability sourced from AutomationBench-AA (Operations and Support), Agentic Customer Interaction removed from capabilities, GDP.pdf added to Long-Context…

Context
Artificial Analysis builds 'Capability Indices' that try to score AI models on real job tasks rather than abstract puzzles: they take the task list behind specific U.S. occupations from the O*NET database, match each task to a that represents it, and weight benchmarks by how often that kind of work actually appears in the job.
In version 1.1, Artificial Analysis added a new agentic-tool-use benchmark to the Finance, Strategy and Ops, Legal, and Healthcare indices, dropped an agentic-customer-interaction benchmark from all of them, and added long- and document-based tests to several. Because the indices are reweighted rather than just extended, a model's rank in a given profession's index can shift even if its own scores haven't changed.
Under the new weighting, Artificial Analysis reports Claude Fable 5.1 still leads all six indices, with GPT-6 Astra second in Finance and Accounting, Strategy and Ops, Legal, and Engineering, while models take narrower wins: Kimi K3 in Finance and Accounting, Legal and Economics, DeepSeek V4.1 Flash in Strategy and Ops, and GLM-5.3 in Healthcare and Engineering.
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
- tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
Checking sign-in…
Loading comments…






