← All IntelClip / CybersecurityResults on 41 verified V8 vulnerabilities: crash rate hides everything
From Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd · ≈21:12
Quantifies why the metric choice matters — near-parity on crashes (and ~50% even for weaker models) collapses to 73% / 68% / 0% once full arbitrary code execution is required.
What’s in it
- Quantifies why the metric choice matters — near-parity on crashes (and ~50% even for weaker models) collapses to 73% / 68% / 0% once full arbitrary code execution is required.
Clip transcript
in this. So, we ran this on 41 V8 vulnerabilities. We went and hand vulnerified verified that they were all exploitable. We took actually the leader for the current Chrome security, his name is Sung Hin Lee, verify these for us. And what we found is that if you're purely looking at old benchmarks where our triggering a crash is what you want to do, it's really not a distinguisher among models. GPT and GPT 5.5 and Mythos both achieved 95%. They were able to trigger a vulnerability 39 out of 41 times. Essentially, all the tasks are side. And then if you started to look at lower powered models, things like Gemini, Kimmy, Minimax, GLM, they were still able to succeed about 50% of the time. So, think about this. If you were looking at the old benchmarks, the message would be 50% of the time Kimmy succeeds in hacking, but that's because their definition of hacking was broken. It was simply crashing it. The real question is can they do a full sandbox escape? And this is where we see a distinguishing characteristics. So, if we look at what I'd call arbitrary code execution is really what the elite would do, Mythos was a quite surprising able to do this 73% of the time. So, 30 out of the 41 examples, Mythos was able to do this sort of full control flow hijack. GPT, sorry, the little bar here is wrong. This was 68% of the time, and Gemini and Kimmy were 0% of the time. So, we're starting to see a signal between these models on what they can do. Little bars here are wrong, but the actual numbers are correct.
Comments
Checking sign-in…
Loading comments…