ExploitGym
cybergym.io- Category
- Cybersecurity
- Rank
- No. 584Tools index
- Listed in
- #38 Find AI benchmarks
- Pricing
- Free
- Platform
- web
- Type
- TOOL
- Date
About
Evaluates whether agents can turn a known vulnerability and a triggering input into a working exploit in controlled evaluation environments.
What it can do
Provide benchmark tasks that test AI agents' ability to convert a vulnerability and proof-of-vulnerability input into a working exploit
Known vulnerability and proof-of-vulnerability input → Working exploit achieving code execution
Evaluate exploits for bypassing mitigations like ASLR and sandboxes
AI-generated exploit code → Success/failure assessment against mitigations
Why it made the leaderboard
Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.
Tags
security-benchmarkexploit-generationai-agentsvulnerability-researchllm-evaluationoffensive-securitylinux-kernelv8
Media

Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.