← All IntelClip / Developer ToolsWhat DeepSWE is: 113 original tasks, near-zero repo overlap
From DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve · ≈0:59
“The median task per repository for us is one.”
“So you can see across over a 100 tasks we pull from nearly 100 repositories.”
What’s in it
- Introduces Deepu, a 113-task long-horizon software engineering benchmark
- Explains how original tasks (not scraped PRs) block agent cheating and contamination
- Shows task diversity across ~100 repos and 5 languages vs rivals
Clip transcript
frontier coding benchmark. So um deepu is a long horizon software engineering software engineering benchmark comprised of 113 original software engineering tasks. So this means unlike something like sweet bench pro we didn't scrape this from um existing PRs that have been closed. Um there's a variety of benefits for this. Uh namely one of them is to resist against contamination and agents being able to cheat uh through the course of their rollouts. Um Swebench has or Swebench Pro uh pulls thousands of tasks from only 40 repositories. The median task per repository for us is one. So you can see across over a 100 tasks we pull from nearly 100 repositories. And the language spans across uh Typescript, JavaScript, Python, Rust, and Go. And we have plans to add uh more languages later on. Um
Comments
Sign in to comment.
Loading comments…