benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Frontier image editors differ far more in consistency than in peak capability. Pass rates span 34% to 83% on targeted edits, so the honest cost metric is attempts per success, not price per image. Protocol and judge code are published for independent runs.