Vibeleaderboard
← All Intel
Intel / post

Safety blocks hit defensive cyber work: CyberGym results and costs

Source
ArtificialAnlys
Date
From the Daily Brief

Artificial Analysis ran its CyberGym-E2E-AA evaluation, which asks models to find memory-safety bugs in large codebases, and reported a split result. Several frontier models refused most of the tasks, in some cases more than 85%, because the discovery steps of defensive work look the same as offensive exploitation to a safety filter. Cheaper models that did engage, including GPT-6 Luna and Xiaomi's MiMo-V2.6-Pro, completed roughly 100 bug hunts for about $20. For security teams, this changes how to pick a model. A higher general intelligence score does not help if the model declines the job, so refusal rate on your own defensive tasks belongs in the evaluation alongside accuracy and cost. It also explains why labs are building gated access programs for defenders, since a general filter cannot tell intent from the request alone.

Read the 2026-10-01 Brief →
ArtificialAnlys@ArtificialAnlys

Restricting offensive without blocking defensive is difficult, as finding and proving a vulnerability requires the same steps whether the goal is to exploit it or patch it On CyberGym-E2E-AA, a benchmark that measures cyber defense capabilities from discovery to patching on memory-safety tasks, some frontier intelligence models are safety blocked from responding on 85%+ of tasks. The good news is that some of the most capable models are also the most cost effective. With GPT-6 Luna or MiMo-V2.6-Pro, you can run ~100 bug hunts in a 1M+ line codebase for ~$20 - up to 100x cheaper per task than the next most capable model Grok 4.7.

More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…