Lab Rank #01 · 82.2 / 100 ·
Gemini
Flagship models; read text, images, audio and video
3.8
3.7
3.6
Lab Rank #01 · 82.2 / 100 ·
Flagship models; read text, images, audio and video
3.8
3.7
3.6
Lab Rank #02 · 82.1 / 100 ·
Flagship chat models, from cheap Nano to costly Pro
6.1
6
5.6
Lab Rank #03 · 77.6 / 100 ·
Flagship chat models, from small to very large
3.8
3.7
3.5
Lab Rank #04 · 72.6 / 100 ·
Flagship model; the lab's top Intelligence Index score
5.5
5
4.8
Lab Rank #05 · 68.7 / 100 ·
Flagship multimodal reasoning for agents and coding
1.3
1.2
30B
Lab Rank #06 · 68.0 / 100 ·
Flagship reasoning models for coding and agentic work
4.7
4.6
4.5
Lab Rank #07 · 67.4 / 100 ·
V2.6
V2.5
V2
Lab Rank #08 · 63.3 / 100 ·
0905
Lab Rank #09 · 63.0 / 100 ·
Flagship chat models, with cheap Flash and larger Pro
V4.1
V4
V3.2
Lab Rank #10 · 62.2 / 100 ·
5.3
5.2
5.1
Lab Rank #11 · 57.9 / 100 ·
Language models for coding and agentic work
01
Lab Rank #12 · 52.7 / 100 ·
Flagship chat models in Small, Medium and Large
4
3
2603
Lab Rank #13 · 52.1 / 100 ·
Small model for reasoning with limited memory
4
Lab Rank #14 · 51.2 / 100 ·
Open reasoning models, from Nano to Ultra
3.5
3
Lab Rank #15 · 50.5 / 100 ·
Agentic model with configurable reasoning levels
Lab Rank #16 · 45.6 / 100 ·
Flagship chat models for RAG, tools and enterprise agents
A
R7b
R
Benchmark standing uses independently published results that share the same protocol, scope, unit, and evidence date. The best comparable release represents each lab in a benchmark run.
A lab needs at least five comparable runs and three canonical releases to rank. Missing evidence stays missing; scores are not inherited across a model family. This measures represented Atlas evidence, not company quality or importance.