Vibeleaderboard
← All Intel
Intel / post

Artificial Analysis launches Cyber Index for AI cyber defense evaluation

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys

Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent…

Read the full post on X
Key takeaways · AI-distilled
  • The index combines three benchmarks: CWE-Bench-AA from Collinear (120 held-out audit-and-patch tasks spanning all OWASP Top 10 2025 categories), DeepsecBench-AA from Vercel (discovery) and CyberGym-E2E-AA from Berkeley RDI (end to end).
  • DeepsecBench-AA gives the a codebase and a budget and scores it against human security reviewers' golden findings: real findings earn credit, and benign code flagged as vulnerable is penalized.
  • CyberGym-E2E-AA requires the agent to find a memory-safety bug, write a proof of concept that triggers the crash, then patch the code so the crash no longer reproduces.
  • Artificial Analysis reports Grok 4.7 (xhigh) and MiMo-V2.6-Pro leading at 56, followed by GPT-6 Luna (max) at 53, GLM-5.3-Flash at 50 and Muse Spark 1.3 (xhigh) at 44.
  • Per Artificial Analysis, safety refusals leave several frontier models 19 to 31 points behind the leaders. Most of the gap is CyberGym-E2E-AA, where GPT-6 Sol and Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 99%.
Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Gives engineers a shared, multi- measure of how well models find, patch, and exploit-test vulnerabilities, covering discovery, patching, and end-to-end tasks. Useful for choosing models for security agents.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…