DeepSWE
deepswe.datacurve.ai- Category
- Developer Tools
- Rank
- No. 2127Tools index
Previous survey · No. 2132 ·
- Listed in
- #3 Find AI benchmarks
- Pricing
- Free
- Type
- TOOL
- Use case
- Model & Agent Evaluation
- GitHub
- 1.8k stars
- Date
About
A coding benchmark from Datacurve built to resist contamination: its 113 tasks are written from scratch rather than adapted from existing commits or pull requests, span 91 repositories and five languages, and are graded by hand-written verifiers that test behaviour rather than implementation. Every model runs on the same mini-swe-agent harness, and reference solutions are roughly 5.5x more code than comparable benchmarks. The leaderboard covers 24 frontier models, with the task data and harness code published.
Why it made the leaderboard
A coding benchmark whose 113 tasks were written from scratch rather than mined from merged pull requests — the design choice that keeps them out of model training data, which is the contamination steadily eroding SWE-bench's signal. Every model runs on the same mini-swe-agent harness and the task data is published, so a score is reproducible rather than a vendor claim.
Intel on DeepSWE
- GPT-6 Astra lands with 75.2% DeepSWE and desktop computer-use scores
- GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
- GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
- GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
- DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Tags
Media

Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.