WeirdML
htihle.github.io- Category
- Developer Tools
- Type
- TOOL
- Date
About
Tests whether models can write working PyTorch code for unfamiliar, well-specified machine-learning problems under a fixed compute budget. Scores come from a non-agentic setup, so agent results are not comparable.
Why it made the leaderboard
Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.
Tags
benchmarkevaluation
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.