Vibeleaderboard
Index / article

Legacy-Bench

factory.ai
Visit factory.ai
Category
Developer Tools
Type
ARTICLE
Added
Jul 28, 2026

About

Legacy-Bench is a benchmark developed by Factory.ai to evaluate how well frontier AI coding agents can understand, maintain, and modify legacy software systems written in languages like COBOL, Fortran, and Java 7. It targets the kind of mission-critical infrastructure — financial settlement systems, telecom routing, and similar — that still runs on decades-old code, testing whether AI agents can handle real-world legacy engineering challenges rather than modern greenfield codebases.

What it can do

  • Benchmark AI coding agents on their ability to understand, maintain, and modify legacy codebases

    Legacy source code (COBOL, Fortran, Java 7)Performance evaluation results

  • Simulate real-world legacy engineering tasks such as financial settlement and telecom routing systems

    Legacy system code and scenariosAgent task outcomes

Why it made the leaderboard

If you're pointing coding agents at decades-old COBOL, Fortran or Java 7 systems, this is one of the few evals that measures that specific competence instead of greenfield modern-stack tasks, giving you a reference point for whether frontier agents can be trusted on maintenance work in mission-critical legacy code.

Tags

benchmarklegacy-softwarecobolfortranjavaai-agentscode-evaluationsoftware-engineering

Media

Legacy-Bench

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.