Vibeleaderboard

Model Atlas

Google DeepMind

Google

Lab Rank #01 · 82.2 / 100 ·

Gemini

Flagship models; read text, images, audio and video

3.8

3.7

3.6

OpenAI

OpenAI

Lab Rank #02 · 82.1 / 100 ·

GPT

Flagship chat models, from cheap Nano to costly Pro

6.1

6

5.6

Qwen

Qwen

Lab Rank #03 · 77.6 / 100 ·

Qwen

Flagship chat models, from small to very large

3.8

3.7

3.5

Anthropic

Anthropic

Lab Rank #04 · 72.6 / 100 ·

Opus

Flagship model; the lab's top Intelligence Index score

5.5

5

4.8

Meta

Meta

Lab Rank #05 · 68.7 / 100 ·

Muse

Flagship multimodal reasoning for agents and coding

1.3

1.2

30B

xAI

xAI

Lab Rank #06 · 68.0 / 100 ·

Grok

Flagship reasoning models for coding and agentic work

4.7

4.6

4.5

Xiaomi

Xiaomi

Lab Rank #07 · 67.4 / 100 ·

Mimo

V2.6

V2.5

V2

Moonshot AI

Moonshot AI

Lab Rank #08 · 63.3 / 100 ·

Kimi

 

0905

DeepSeek

DeepSeek

Lab Rank #09 · 63.0 / 100 ·

DeepSeek

Flagship chat models, with cheap Flash and larger Pro

V4.1

V4

V3.2

Z.ai

Z.ai

Lab Rank #10 · 62.2 / 100 ·

GLM

5.3

5.2

5.1

MiniMax

MiniMax

Lab Rank #11 · 57.9 / 100 ·

Minimax

Language models for coding and agentic work

 

01

Mistral AI

Mistral AI

Lab Rank #12 · 52.7 / 100 ·

Mistral

Flagship chat models in Small, Medium and Large

4

3

2603

Microsoft

Microsoft

Lab Rank #13 · 52.1 / 100 ·

PHI

Small model for reasoning with limited memory

4

NVIDIA

NVIDIA

Lab Rank #14 · 51.2 / 100 ·

Nemotron

Open reasoning models, from Nano to Ultra

3.5

3

Tencent

Tencent

Lab Rank #15 · 50.5 / 100 ·

HY3

Agentic model with configurable reasoning levels

Cohere

Cohere

Lab Rank #16 · 45.6 / 100 ·

Command

Flagship chat models for RAG, tools and enterprise agents

A

R7b

R

Lab Rank methodology+

Benchmark standing uses independently published results that share the same protocol, scope, unit, and evidence date. The best comparable release represents each lab in a benchmark run.

A lab needs at least five comparable runs and three canonical releases to rank. Missing evidence stays missing; scores are not inherited across a model family. This measures represented Atlas evidence, not company quality or importance.