SEC-Bench Pro
sec-bench.github.io- Category
- Cybersecurity
- Rank
- No. 1272Tools index
- Listed in
- #37 Find AI benchmarks
- Platform
- web
- Type
- TOOL
- Date
About
Evaluates long-horizon security bug hunting in complex systems, including browser engines and the Linux kernel, with reproducible validation.
What it can do
Evaluate AI agents on long-horizon security vulnerability-hunting tasks
AI agent → Task success results
Generate proof-of-concept exploits for real-world software targets
Target codebase (e.g., V8, Firefox, Linux) → Proof-of-concept exploit
Why it made the leaderboard
Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.
Tags
benchmarkllm-evaluationsecurityvulnerability-researchai-agentspoc-generation
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.