Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank
No. 1152Tools index
Pricing
Open Source
Type
TOOL
Added
Aug 19, 2026

About

An OpenAI benchmark measuring how well agents do machine-learning engineering, using 75 Kaggle competitions as the task set and Kaggle's own medal thresholds as the bar. Agents get a fixed budget — 24 hours, 36 vCPUs, 440GB RAM and a single 24GB A10 — and are scored on the share of competitions where they take any medal, run over at least three seeds and reported with standard error. A 22-competition low-complexity split exists for cheaper runs.

Why it made the leaderboard

It measures agents doing machine-learning engineering end to end — 75 Kaggle competitions, scored by whether the agent's submission would have taken a medal. Every run gets the same 24 hours, 36 vCPUs and one A10, and results are averaged over at least three seeds with standard error, so a headline number means something.

Tags

benchmarkml-engineeringagentskaggleevaluationopenai

Tech Stack

Python

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.