Vibeleaderboard
Index / tool

Humanity's Last Exam

lastexam.ai
Visit lastexam.ai
Category
Developer Tools
Rank
No. 1360Tools index
Pricing
Free
Type
TOOL
Added
Aug 19, 2026

About

A closed-ended academic benchmark built to outlast frontier models: 2,500 questions across more than 100 subjects, written by close to 1,000 expert contributors at over 500 institutions in 50 countries, and published in Nature in January 2026. Frontier scores stay low — Gemini 3 Pro at 38.3%, GPT-5 at 25.3% — and a private held-out set is kept alongside the public HuggingFace dataset to detect overfitting to the released questions.

Why it made the leaderboard

2,500 questions written by close to a thousand experts across more than a hundred subjects, built so frontier models cannot saturate it — the leaders still sit under 40%. A private held-out set is kept beside the public dataset, which is how you tell a real capability gain from a model that has simply seen the questions.

Tags

benchmarkreasoningfrontier-modelsevaluationleaderboardacademic

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.