
contamination inflates reported coding scores. CGMIA catches leaked samples that perplexity-only detectors miss, giving practitioners a more reliable way to judge whether a leaderboard score reflects real capability.
“Code generation benchmarks are widely used to evaluate Large Language Models (LLMs), but benchmark data leakage into training sets can inflate performance and undermine evaluation validity.”
“Experiments on eight code generation benchmarks show that CGMIA outperforms eight existing membership inference methods in most cases.”
articleHow effective are traditional test criteria at detecting bugs in large language models generated code?Asma Hamidi, Michael Konstantinou, Renzo Degiovanni, Mike Papadakis
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
articleMemorization Diagnostics for Code LLMs Should be Scale-AwarePrateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir\'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend\'e F. Bissyand\'eChecking sign-in…
Loading comments…