Vibeleaderboard
Index / tool
Visit sec-bench.github.io
Category
Cybersecurity
Rank
No. 1272Tools index
Listed in
#37 Find AI benchmarks
Platform
web
Type
TOOL
Date

About

Evaluates long-horizon security bug hunting in complex systems, including browser engines and the Linux kernel, with reproducible validation.

What it can do

  • Evaluate AI agents on long-horizon security vulnerability-hunting tasks

    AI agentTask success results

  • Generate proof-of-concept exploits for real-world software targets

    Target codebase (e.g., V8, Firefox, Linux)Proof-of-concept exploit

Why it made the leaderboard

Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.

Tags

benchmarkllm-evaluationsecurityvulnerability-researchai-agentspoc-generation

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.