Vibeleaderboard
← All Intel
Intel / article

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

Source
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
Author
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
Date
Key takeaways · AI-distilled
  • Three Qwen3 judges did not open with a verdict on 12% to 49% of pairs, versus under 3% for Llama-3.1-8B and Phi-3.5-mini. Forcing a first-token read on those pairs returns whichever response was shown first.
  • On the 924 pairs where a judge did not commit, the forced read flipped on 89.7% when responses were swapped, against 47.5% when the verdict was read after generation.
  • A smaller failure remains even when a judge leads with a verdict: it sometimes opens with one letter and reasons its way to the other, on 0 to 5.5% of pairs, at a rate uncorrelated with compliance.
  • The authors recommend reporting how often a judge leads with a verdict token, which costs one forward pass and no labels, next to any position-bias figure, and treating forced-read bias numbers as upper bounds.
Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
Why it matters

If your scores judges from first-token logits, measured position bias is inflated by about 42 points. Generate the verdict before auditing a judge.

Recommended reads
Comments

Checking sign-in…

Loading comments…