← Back to Vibers

Builder
SWE-bench
1 Tool
SWE-bench is a benchmark, created by researchers at Princeton and Stanford and introduced at ICLR 2024, that evaluates language models on their ability to resolve real GitHub issues by generating working code patches. Built from roughly 2,294 task instances drawn from pull requests across a dozen popular Python repositories, it has become a standard yardstick for AI coding-agent performance, with variants including SWE-bench Lite and SWE-bench Verified.