Vibeleaderboard
← All Intel
Intel / article

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

Source
Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
Author
Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
Date
Key takeaways · AI-distilled
  • Run unmodified on two code-judging benchmarks, the MARCH verification framework called both candidate solutions equally good in 78% to 95% of comparisons, reaching 4.4% accuracy where the same model asked directly scored 43.7%.
  • The authors argue claim-by-claim verification needs evidence that is independent of the answer and differs between the two candidates, a condition retrieved documents satisfy but code judging does not.
  • Neither easier problems nor a larger judge model changed the result; the authors explain the collapse with two label-free measurements taken from the pipeline's own logs.
  • Gating on a label-free signal from the pipeline's own logs let the judge decline comparisons it could not ground, raising accuracy from 20.7% to 36.9% while still answering half of all comparisons.
Terms in this piece · Glossary
  • multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
Why it matters

MARCH, a multi-agent code-judging method, calls ties on 78-95% of comparisons and reaches only 4.4% accuracy where directly asking the model scores 43.7%, showing multi-agent verification breaks when evidence can't differ between the candidates being judged.

Recommended reads
Comments

Checking sign-in…

Loading comments…