Vibeleaderboard
Index / tool

ExploitGym

cybergym.io
Visit cybergym.io
Category
Cybersecurity
Rank
Listed in
#38 Find AI benchmarks
Pricing
Free
Platform
web
Type
TOOL
Date

About

Evaluates whether agents can turn a known vulnerability and a triggering input into a working exploit in controlled evaluation environments.

What it can do

  • Provide benchmark tasks that test AI agents' ability to convert a vulnerability and proof-of-vulnerability input into a working exploit

    Known vulnerability and proof-of-vulnerability inputWorking exploit achieving code execution

  • Evaluate exploits for bypassing mitigations like ASLR and sandboxes

    AI-generated exploit codeSuccess/failure assessment against mitigations

Why it made the leaderboard

Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.

Tags

security-benchmarkexploit-generationai-agentsvulnerability-researchllm-evaluationoffensive-securitylinux-kernelv8

Media

ExploitGym

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.