Vibeleaderboard
← Back to Vibers
SWE-bench
Builder

SWE-bench

1 Tool

SWE-bench is a benchmark, created by researchers at Princeton and Stanford and introduced at ICLR 2024, that evaluates language models on their ability to resolve real GitHub issues by generating working code patches. Built from roughly 2,294 task instances drawn from pull requests across a dozen popular Python repositories, it has become a standard yardstick for AI coding-agent performance, with variants including SWE-bench Lite and SWE-bench Verified.

Tools

SWE-bench(swebench.com)

Benchmark testing whether language-model agents can resolve real GitHub issues drawn from production codebases.

Developer ToolsLLM Eval5.8kMITbuilt by swe-bench