Vibeleaderboard
← All Intel
Intel / post

Artificial Analysis Adds Per-Effort-Level Model Comparison Pages

Source
ArtificialAnlys
Date
From the Daily Brief

Artificial Analysis has added model-release pages that compare every available reasoning-effort setting for a model across intelligence, cost per task, output speed, and latency. Frontier models can now ship with as many as six effort levels, even though the underlying weights stay the same. More effort can improve results, but it also changes response time and spend, so a model name alone no longer defines the product a team is buying. The new pages place those configurations on the same charts and break out performance across finance, legal, healthcare, strategy, engineering, and economics tasks. Artificial Analysis reports that GPT-6 Astra spans 46 to 53 on its Intelligence Index as time per task rises from 1.6 to 8.2 minutes across effort settings. Those figures come from the publisher's own benchmark suite, so they should guide a workload test rather than replace one. The useful change is visibility: teams can choose an effort level from measured tradeoffs instead of assuming that the highest setting is always the best operational default.

Read the 2026-09-09 Brief →
ArtificialAnlys@ArtificialAnlys

You can now easily compare the intelligence, cost, and speed on different effort levels for your preferred models with the new Artificial Analysis Model Release pages How a model performs across intelligence and cost is heavily influenced by its configured reasoning and effort level. Frontier models are now being released with up to six different effort levels, meaning the same model weights can produce significantly different performance and cost profiles. Model Release pages make it easy to compare these configurations across all of our standard charts in one place Each release page features: ➤ Intelligence, cost per task, output speed, and latency across every effort variant ➤ Side by side comparison of effort levels against those of similar models ➤ Artificial Analysis Capability Index scores on each effort level across Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics For example, on Intelligence vs Time per Task, GPT-6 Astra ranges from 46–53 on the Artificial Analysis Intelligence Index and 1.6–8.2 minutes per task, while Claude Fable 5.1 spans a similar intelligence range (47–53) on a wider time per task range (4.2–12.2…

Read the full post on X

Context

Artificial Analysis measures models on a combined Intelligence Index score and tracks what each task costs to run. Its new Model Release pages group those measurements by effort level, the configurable amount of internal reasoning a model does before answering, since the same underlying model weights can score very differently and cost very differently depending on that setting.

As a concrete example from the post, GPT-6 Astra scored 46 to 53 on the Intelligence Index while taking 1.6 to 8.2 minutes per task across its effort settings, and Claude Fable 5.1 scored a similar 47 to 53 but across a wider 4.2 to 12.2 minutes per task. That means comparing two models by intelligence score alone can hide a large difference in how long a task takes to complete, which the new pages are meant to surface.

Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Key quotes

“Frontier models are now being released with up to six different effort levels, meaning the same model weights can produce significantly different performance and cost profiles.”

More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…