DeepSWE
deepswe.datacurve.ai- Category
- Developer Tools
- Rank
- No. 1544Tools index
- Pricing
- Free
- Type
- TOOL
- GitHub
- 1.4k stars
- Added
- Aug 19, 2026
About
A coding benchmark from Datacurve built to resist contamination: its 113 tasks are written from scratch rather than adapted from existing commits or pull requests, span 91 repositories and five languages, and are graded by hand-written verifiers that test behaviour rather than implementation. Every model runs on the same mini-swe-agent harness, and reference solutions are roughly 5.5x more code than comparable benchmarks. The leaderboard covers 24 frontier models, with the task data and harness code published.
Why it made the leaderboard
A coding benchmark whose 113 tasks were written from scratch rather than mined from merged pull requests — the design choice that keeps them out of model training data, which is the contamination steadily eroding SWE-bench's signal. Every model runs on the same mini-swe-agent harness and the task data is published, so a score is reproducible rather than a vendor claim.
Intel on DeepSWE
- DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and RoutingAug 18, 2026
- DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and RoutingAug 18, 2026
- DeepSWE v1.1 hardens against reward hackingAug 11, 2026
- What DeepSWE is: 113 original tasks, near-zero repo overlapAug 11, 2026
- DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and CodingAug 7, 2026
Tags
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.