Researchers checked whether an eight-mechanism methodological agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → for agentic development, from context engineeringDeciding what an agent sees at each step — the discipline that replaced prompt engineering once context windows got large enough to fill badly.Full definition → to graduated autonomy, shows up in practice across 5,435 repositories sampled from the 116,211-repo AIDev dataset.
At least one mechanism appears in 21.7% of the population but in 65.3% of the most visible repositories, measured by stars. Repositories combining several mechanisms are rare.
Coding rule files from 150 repositories showed they almost always guide the AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → and state norms, while shared persistent knowledge, executable specifications, structured consultation and graduated autonomy appear only in a minority.
The authors conclude the harness is observable but mostly as isolated mechanisms, not the integrated system the framework proposes.
Terms in this piece · Glossary
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
context engineering — Deciding what an agent sees at each step — the discipline that replaced prompt engineering once context windows got large enough to fill badly.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters
Shows how teams actually configure coding agents in real repositories, including rule and context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → files, specs and decision records. It helps you compare your setup against observed adoption of context engineering and graduated autonomy.