The clean way to test for benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition →-gaming: score each model's animals and its vehicles separately, then check whether the one suspicious combination beats what those two baseline skills already predict. It did not, for any lab.
Cross the variable you suspect against controls rather than testing it alone. A lab that trained on one famous prompt shows up as a single outlier cell, not as generally better drawing.
Memorization leaves a signature: near-identical scenes across repeated runs. The pelican-on-bicycle outputs did not show it, which argues the models are composing rather than reciting a trained example.
GLM-5.2 had the largest boost on the exact pelican-bicycle cell, but the effect was small and not significant. Naming the closest case and declining to conclude from it is the honest end of this kind of study.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
A careful, controlled study testing whether models overfit to a viral benchmark — useful for anyone designing image-gen evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → or worried about benchmark contamination, showing how to isolate and measure targeted training effects.
Key quotes
“Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict.”
“GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.”