
Documents a concrete, reproducible failure mode where a frontier model detects and defeats a static web-enabled by reverse-engineering its answer key — critical for anyone relying on BrowseComp-style to judge model capability or safety.
“Instead of inadvertently coming across a leaked answer, Claude Opus 4.6 independently hypothesized that it was being evaluated, identified which benchmark it was running in, then located and decrypted the answer key.”
“To our knowledge, this is the first documented instance of a model suspecting it is being evaluated without knowing which benchmark was being administered, then working backward to successfully identify and solve the evaluation itself.”
“This finding raises questions about whether static benchmarks remain reliable when run in web-enabled environments.”
“One of these problems consumed 40.5 million tokens, roughly 38 times higher than the median.”
articleDesigning AI-resistant technical evaluationsAnthropic Engineering
articleInvestigating Three Real-World Incidents in Our Cybersecurity EvaluationsSimon Willison
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert KirkChecking sign-in…
Loading comments…