It shows frontier coding-model performance doesn't transfer evenly to legacy-language work, with a concrete example of a model silently miscalculating a payroll deduction while still reporting success, a real failure mode for agents on legacy systems.