CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
Source
arxiv.org
Author
Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, Zaiyuan Wang, Haiying Sun, Ting Su, Chengcheng Wan
Date
Why it matters
Most coding-agent benchmarks test patching. This one measures whether agents can write working static-analysis checkers from a CVE spec, across 21 model-harness setups, exposing long-horizon and security tooling gaps.
Key takeaways · AI-distilled
Checker synthesis asks an AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → to interpret a defect specification, inspect the repository, write analyzer-specific logic and refine it through repeated compile-and-analyze feedback, which existing coding-agent benchmarks rarely assess end to end.
Each task ships vulnerable and fixed revisions, a pinned analysis environment and a checker scaffold, spanning 85 CWEs.
A companion framework, CheckerLab, independently rebuilds every submitted checker and also tracks how the agent used tools, alongside diagnostic contrast, localization and false positives.
Across 21 model-agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → configurations with three repeats each, mean Pass@1 was 32.30% and the best configuration reached 45.33%.
Terms in this piece · Glossary
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.