Vibeleaderboard
← All Intel
Intel / video

The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals

Source
youtube.com
Author
Latent Space
Date
Why it matters

progress has stalled because the is saturated and contaminated, so scores on it no longer separate coding models. Check what replaces it before trusting leaderboard claims.

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Read the source www.youtube.com
More from Latent Space
Recommended reads
Comments

Checking sign-in…

Loading comments…