Benchmark: AI doesn't find bugs unless you tell it what's wrong
Source
lieret
Author
lieret
Date
Key takeaways · AI-distilled
SWE-Sweep tests whether models find and fix bugs nobody has reported yet. It draws on about 4,000 real GitHub bugs across 100 repos in 22 languages, filtered so each one is solvable in that setting.
The best setup the authors tested, Sol 5.6 at xhigh effort, scored only 4.7% at a reported cost of about $7,230. Luna 5.6 at xhigh reached 2.5% for about $224.
Other setups trailed further: Opus 5 at xhigh scored 1.3% for about $5,363, Kimi K3 0.6%, and GPT-5.4 Mini and Gemini 3.5 Flash Lite 0.5% or below. The benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → is MIT-licensed on GitHub under facebookresearch.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Models fix reported bugs well but proactive bug discovery is still near 5% at best, at high cost. This sets realistic expectations for autonomous code-review and bug-hunting agents.