OpenCollab argues multi-agentUsing several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.Full definition → coding evaluations typically assume agents faithfully follow their configured organization, and defines an Adherence metric that measures whether the declared organization is actually realized at runtime.
Changing any single dimension of the configuration shifted Adherence from 47.2% to as high as 97.2%, the authors report, so agents often collaborate quite differently from how they were set up.
The framework runs every configuration on a shared, controlled runtime and records fine-grained event streams, so gains can be attributed to the organization rather than to differences in system components.
A two-coder workflow built on OpenCollab reportedly set new state-of-the-art results against Mini-SWE-AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →, Codex CLI and Claude Code, while its single-agent configuration used the fewest tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → across the evaluated suites. The paper is marked work in progress.
Terms in this piece · Glossary
multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
Why it matters
Questions multi-agent coding gains by measuring whether agents actually follow their assigned organization, so reported improvements can be attributed to the right cause.