← All IntelClip / EducationContamination in SWE-bench Verified via Opus memorization
From Benchmaxxing: The Gap Between Benchmark Scores and Reality · ≈4:41
“contamination is the default outcome unless you are very very good”
“You can give opus the first part of the prompt and it will verbatim spit out the rest.”
“we found very clear evidence that Opus had memorized a lot of Sweetbench”
“as benchmarking consumers, we're just missing that”
What’s in it
- Reveals Opus can verbatim regurgitate SWE-bench Verified prompts and answers
- Shows a method for testing if models memorized specific benchmark repos
- Exposes that model cards report SWE-bench scores without contamination disclosures
Clip transcript
Contamination is often thought of as when labs are explicitly training on the test set and that does happen sometimes but really contamination is the default outcome unless you are very very good. So labs put a lot of effort into holding back this flood of data that's going to contaminate their models. But inevitably if you have public questions and answers on the internet that's going to get memorized to some extent. So SweetBench verified here's an example prompt. You can give opus the first part of the prompt and it will verbatim spit out the rest. It does that with the answers as well. And we actually did an investigation where we compared looking at the repos that Sweepbench verified was built out of. How much has Opus memorized the Sweepbench verified contents versus the rest of the repo? And we found very clear evidence that Opus had memorized a lot of Sweetbench. In the most recent model card, Opus 4.8 talks about its SWE score. It does not disclose this contamination. We as an industry aren't really in the habit of doing those disclosures. And so what that means is that as benchmarking consumers, we're just missing that
Comments
Sign in to comment.
Loading comments…