FrontierCode
cognition.com- Category
- Developer Tools
- Type
- TOOL
- Date
About
Cognition's benchmark of hard, real open-source issues where coding agents must produce mergeable fixes. Epoch AI found too little public information to review it, so treat reported scores cautiously.
Why it made the leaderboard
Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.
Tags
benchmarkevaluation
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.