If you're pointing coding agents at decades-old COBOL, Fortran or Java 7 systems, this is one of the few evals that measures that specific competence instead of greenfield modern-stack tasks, giving you a reference point for whether frontier agents can be trusted on maintenance work in mission-critical legacy code.
Legacy-Bench is a benchmark developed by Factory.ai to evaluate how well frontier AI coding agents can understand, maintain, and modify legacy software systems written in languages like COBOL, Fortran, and Java 7.
It targets the kind of mission-critical infrastructure — financial settlement systems, telecom routing, and similar — that still runs on decades-old code, testing whether AI agents can handle real-world legacy engineering challenges rather than modern greenfield codebases.
Transcript
A new benchmark designed to measure frontier AI agent capabilities on legacy software engineering tasks spanning COBOL, Fortran, Java 7, and more.