
AgentRuleBench
github.com/tommkruix/agentrulebench- Category
- Developer Tools
- Rank
- No. 2273Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 3 stars
- Date
About
AgentRuleBench is a reproducible, pre-registered benchmark testing whether AI coding agents drift from architectural import-boundary rules written in prose files like CLAUDE.md, AGENTS.md, or GEMINI.md, when no automated lint rule enforces them. Across a pilot spanning three vendors' agents and four conditions, from an unguarded control to a run-lint-and-fix setup, the agents did not violate the measured rule, a null result against the common assumption that prose conventions need deterministic lint enforcement.
Why it made the leaderboard
It tests, and largely refutes, the common assumption that AI coding agents need deterministic lint rules to hold architectural boundaries described only in CLAUDE.md/AGENTS.md-style prose.
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.