Vibeleaderboard
Index / tool

SWE-rebench

swe-rebench.com
Visit swe-rebench.com
Category
Developer Tools
Rank
No. 2768Tools index

Previous survey · No. 2623 ·

Listed in
#4 Find AI benchmarks
Pricing
Free
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

A continuously refreshed software-engineering benchmark that draws new tasks from recent repository activity, so models are scored on problems published after their training cutoff. The current window covers 111 problems across 65 repositories, reported as resolved rate and pass@5 alongside cost and token use per problem, with rows flagged for potential contamination, external scaffolding, or falling outside the evaluation window. Problems ship as a HuggingFace dataset with a Docker image per problem.

Why it made the leaderboard

It re-draws its tasks from recent repository activity, so models are scored on problems that postdate their training cutoff — the one structural answer to benchmark contamination. Each row reports cost and tokens per problem and carries flags for suspected contamination or external scaffolding, and the problems ship as a public dataset with a Docker image each.

Tags

benchmarkswecoding-agentscontaminationleaderboardevaluation

Media

SWE-rebench

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.