Vibeleaderboard
← All Intel
Intel / article

Benchmark: AI doesn't find bugs unless you tell it what's wrong

Source
lieret
Author
lieret
Date
Key takeaways · AI-distilled
  • SWE-Sweep tests whether models find and fix bugs nobody has reported yet. It draws on about 4,000 real GitHub bugs across 100 repos in 22 languages, filtered so each one is solvable in that setting.
  • The best setup the authors tested, Sol 5.6 at xhigh effort, scored only 4.7% at a reported cost of about $7,230. Luna 5.6 at xhigh reached 2.5% for about $224.
  • Other setups trailed further: Opus 5 at xhigh scored 1.3% for about $5,363, Kimi K3 0.6%, and GPT-5.4 Mini and Gemini 3.5 Flash Lite 0.5% or below. The is MIT-licensed on GitHub under facebookresearch.
Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

Models fix reported bugs well but proactive bug discovery is still near 5% at best, at high cost. This sets realistic expectations for autonomous code-review and bug-hunting agents.

Read the source swesweep.com
Recommended reads
Comments

Checking sign-in…

Loading comments…