Vibeleaderboard
← All Intel
Intel / article

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

Source
Leonardo Ferreira, Gardenia Liu, Kaden Zheng
Author
Leonardo Ferreira, Gardenia Liu, Kaden Zheng
Date
Key takeaways · AI-distilled
  • The study spans 23 small models from eleven vendor families, five tasks and more than 5,500 debate and control runs, pairing every debate setup with a majority-vote control given the same generation budget.
  • Debate beat single- inference by 3 to 7 points on tasks with headroom, but at matched budgets it tied or lost to self-consistency sampling while using 1.6x the wall-clock time and 3.4x the tokens.
  • Persona prompting lowered accuracy. A dose-response test over each model's persona combinations points to a persona tax, not a diversity tax: redundant personas hurt most, and maximally diverse teams recovered part of the loss.
  • Mixed-model teams lost to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate's benefit came from the first exchange of answers.
  • The authors flag a measurement hazard: debate transcripts silently overflowed serving windows, and correcting that alone moved their debate-versus-sampling gap from -1.8 points to parity.
Terms in this piece · Glossary
  • multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
  • open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters

If you are adding debate or persona agents, budget-matched majority voting often matches them at far lower cost. This tests the assumption that agent diversity drives gains.

Recommended reads
Comments

Checking sign-in…

Loading comments…