AsynCodeBench scores multi-agentUsing several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.Full definition → coding on collaboration itself: each of its 19 tasks from real repositories carries an explicit dependency graph, 52 cross-AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → dependencies in total, checked by executable Dependency Checkers.
The authors propose two metrics: Asynchronous Dependency Pass Rate, how many cross-agent dependencies end up satisfied, and Dependency Resolution Step, when each dependency is first satisfied during the run.
Across model families, scales and generations, better coding scores did not reliably mean better collaboration, and task-level results could diverge sharply from dependency-level measures, so task pass rates alone can hide coordination failures.
Trajectory analysis found coordination tends to succeed in bursts, many dependencies resolving within a short stretch of the run, a pattern the authors call a hopping window, rather than improving gradually.
Terms in this piece · Glossary
multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Multi-agent coding is usually scored by task pass rate, which conflates individual coding agent skillA reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.Full definition → with coordination. This benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → measures cross-agent dependency satisfaction directly, letting you evaluate orchestration separately.