When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
Source
Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
Author
Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
Date
Key takeaways · AI-distilled
Run unmodified on two code-judging benchmarks, the MARCH multi-agentUsing several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.Full definition → verification framework called both candidate solutions equally good in 78% to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly scored 43.7%.
The authors argue claim-by-claim verification needs evidence that is independent of the answer and differs between the two candidates, a condition retrieved documents satisfy but code judging does not.
Neither easier problems nor a larger judge model changed the result; the authors explain the collapse with two label-free measurements taken from the pipeline's own logs.
Gating on a label-free signal from the pipeline's own logs let the judge decline comparisons it could not ground, raising accuracy from 20.7% to 36.9% while still answering half of all comparisons.
Terms in this piece · Glossary
multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
Why it matters
MARCH, a multi-agent code-judging method, calls ties on 78-95% of comparisons and reaches only 4.4% accuracy where directly asking the model scores 43.7%, showing multi-agent verification breaks when evidence can't differ between the candidates being judged.