benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Signed by working mathematicians, it argues AI labs treating landmark proofs as benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → trophies threatens the verification, attribution, and teaching norms that make mathematical progress durable, a pattern likely to recur in other fields.