Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank
No. 2273Tools index
Pricing
Open Source
Type
TOOL
GitHub
3 stars
Date

About

AgentRuleBench is a reproducible, pre-registered benchmark testing whether AI coding agents drift from architectural import-boundary rules written in prose files like CLAUDE.md, AGENTS.md, or GEMINI.md, when no automated lint rule enforces them. Across a pilot spanning three vendors' agents and four conditions, from an unguarded control to a run-lint-and-fix setup, the agents did not violate the measured rule, a null result against the common assumption that prose conventions need deterministic lint enforcement.

Why it made the leaderboard

It tests, and largely refutes, the common assumption that AI coding agents need deterministic lint rules to hold architectural boundaries described only in CLAUDE.md/AGENTS.md-style prose.

Tags

benchmarkai-coding-agentsllm-evalarchitecture-rulesnull-result

Tech Stack

Node.js

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.