Vibeleaderboard
← All Intel
Intel / post

From SWE-bench to ProgramBench: The Next Generation of Coding Evals

Source
Stanford AI Lab
Date
Stanford AI Lab@StanfordAILab
Thread · 2 parts

Hear from @jyangballin on ProgramBench and the lineage of AI coding benchmarks! In conversation with @vincentsunnchen https://t.co/VwIbAIsLR3

@jyangballin @vincentsunnchen https://t.co/P8vGVpGyc5

Terms in this piece · Glossary
  • SWE-benchThe standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
  • AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters

Gives coding- builders on where is heading past and what gaps ProgramBench is designed to close.

More from Stanford AI Lab
Recommended reads
Comments

Checking sign-in…

Loading comments…