Vibeleaderboard
Index / tool
Visit deepswe.datacurve.ai
Category
Developer Tools
Rank
No. 1544Tools index
Pricing
Free
Type
TOOL
Added
Aug 19, 2026

About

A coding benchmark from Datacurve built to resist contamination: its 113 tasks are written from scratch rather than adapted from existing commits or pull requests, span 91 repositories and five languages, and are graded by hand-written verifiers that test behaviour rather than implementation. Every model runs on the same mini-swe-agent harness, and reference solutions are roughly 5.5x more code than comparable benchmarks. The leaderboard covers 24 frontier models, with the task data and harness code published.

Why it made the leaderboard

A coding benchmark whose 113 tasks were written from scratch rather than mined from merged pull requests — the design choice that keeps them out of model training data, which is the contamination steadily eroding SWE-bench's signal. Every model runs on the same mini-swe-agent harness and the task data is published, so a score is reproducible rather than a vendor claim.

Intel on DeepSWE

More in Intel

Tags

benchmarkcoding-agentsevaluationleaderboardswecontamination

Media

DeepSWE

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.