Vibeleaderboard
← All Intel
Intel / article

AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering

Source
Kaituo Zhang, Zhen Xiong, Zhimeng Jiang, Mingyu Zhong, Zhouyuan Yuan, Zhecheng Li, Bowen Lin, Chia-Yuan Chang, Mingzhi Hu, Huazheng Wang, Ying Lin
Author
Kaituo Zhang, Zhen Xiong, Zhimeng Jiang, Mingyu Zhong, Zhouyuan Yuan, Zhecheng Li, Bowen Lin, Chia-Yuan Chang, Mingzhi Hu, Huazheng Wang, Ying Lin
Date
Key takeaways · AI-distilled
  • AsynCodeBench scores coding on collaboration itself: each of its 19 tasks from real repositories carries an explicit dependency graph, 52 cross- dependencies in total, checked by executable Dependency Checkers.
  • The authors propose two metrics: Asynchronous Dependency Pass Rate, how many cross-agent dependencies end up satisfied, and Dependency Resolution Step, when each dependency is first satisfied during the run.
  • Across model families, scales and generations, better coding scores did not reliably mean better collaboration, and task-level results could diverge sharply from dependency-level measures, so task pass rates alone can hide coordination failures.
  • Trajectory analysis found coordination tends to succeed in bursts, many dependencies resolving within a short stretch of the run, a pattern the authors call a hopping window, rather than improving gradually.
Terms in this piece · Glossary
  • multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

Multi-agent coding is usually scored by task pass rate, which conflates individual coding with coordination. This measures cross-agent dependency satisfaction directly, letting you evaluate orchestration separately.

Recommended reads
Comments

Checking sign-in…

Loading comments…