Artificial Analysis launches Cyber Index for AI cyber defense evaluation
- Source
- Artificial Analysis
- Date

Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent…

- The index combines three benchmarks: CWE-Bench-AA from Collinear (120 held-out audit-and-patch tasks spanning all OWASP Top 10 2025 categories), DeepsecBench-AA from Vercel (discovery) and CyberGym-E2E-AA from Berkeley RDI (end to end).
- DeepsecBench-AA gives the a codebase and a budget and scores it against human security reviewers' golden findings: real findings earn credit, and benign code flagged as vulnerable is penalized.
- CyberGym-E2E-AA requires the agent to find a memory-safety bug, write a proof of concept that triggers the crash, then patch the code so the crash no longer reproduces.
- Artificial Analysis reports Grok 4.7 (xhigh) and MiMo-V2.6-Pro leading at 56, followed by GPT-6 Luna (max) at 53, GLM-5.3-Flash at 50 and Muse Spark 1.3 (xhigh) at 44.
- Per Artificial Analysis, safety refusals leave several frontier models 19 to 31 points behind the leaders. Most of the gap is CyberGym-E2E-AA, where GPT-6 Sol and Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 99%.
Gives engineers a shared, multi- measure of how well models find, patch, and exploit-test vulnerabilities, covering discovery, patching, and end-to-end tasks. Useful for choosing models for security agents.
articleOpenAI’s accidental cyberattack against Hugging Face is science fiction that happenedSimon Willison
videoTraining Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging FaceAI Engineer
postAnnouncing the Artificial Analysis Search Index, benchmarking how search API…Artificial Analysis
Checking sign-in…
Loading comments…



