
MLE-bench
github.com/openai/mle-bench- Category
- Developer Tools
- Rank
- No. 2163Tools index
- Listed in
- #25 Find AI benchmarks
- Pricing
- Open Source
- Type
- TOOL
- Use case
- Model & Agent Evaluation
- GitHub
- 1.8k stars
- Date
About
An OpenAI benchmark measuring how well agents do machine-learning engineering, using 75 Kaggle competitions as the task set and Kaggle's own medal thresholds as the bar. Agents get a fixed budget — 24 hours, 36 vCPUs, 440GB RAM and a single 24GB A10 — and are scored on the share of competitions where they take any medal, run over at least three seeds and reported with standard error. A 22-competition low-complexity split exists for cheaper runs.
Why it made the leaderboard
It measures agents doing machine-learning engineering end to end — 75 Kaggle competitions, scored by whether the agent's submission would have taken a medal. Every run gets the same 24 hours, 36 vCPUs and one A10, and results are averaged over at least three seeds with standard error, so a headline number means something.
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.