Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank
No. 2163Tools index
Listed in
#25 Find AI benchmarks
Pricing
Open Source
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

An OpenAI benchmark measuring how well agents do machine-learning engineering, using 75 Kaggle competitions as the task set and Kaggle's own medal thresholds as the bar. Agents get a fixed budget — 24 hours, 36 vCPUs, 440GB RAM and a single 24GB A10 — and are scored on the share of competitions where they take any medal, run over at least three seeds and reported with standard error. A 22-competition low-complexity split exists for cheaper runs.

Why it made the leaderboard

It measures agents doing machine-learning engineering end to end — 75 Kaggle competitions, scored by whether the agent's submission would have taken a medal. Every run gets the same 24 hours, 36 vCPUs and one A10, and results are averaged over at least three seeds with standard error, so a headline number means something.

Tags

benchmarkml-engineeringagentskaggleevaluationopenai

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.