benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
The benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → shows general coding ability doesn't transfer to legacy languages like COBOL or Fortran, with scores ranging from 60% to 23% across frontier models and a documented case of a model silently corrupting a payroll calculation while still passing most tests, a concrete reason to distrust plausible-looking output on legacy migrations.