The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating
Source
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
Author
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
Date
Key takeaways · AI-distilled
Three Qwen3 judges did not open with a verdict tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → on 12% to 49% of pairs, versus under 3% for Llama-3.1-8B and Phi-3.5-mini. Forcing a first-token read on those pairs returns whichever response was shown first.
On the 924 pairs where a judge did not commit, the forced read flipped on 89.7% when responses were swapped, against 47.5% when the verdict was read after generation.
A smaller failure remains even when a judge leads with a verdict: it sometimes opens with one letter and reasons its way to the other, on 0 to 5.5% of pairs, at a rate uncorrelated with compliance.
The authors recommend reporting how often a judge leads with a verdict token, which costs one forward pass and no labels, next to any position-bias figure, and treating forced-read bias numbers as upper bounds.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
Why it matters
If your evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition →agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → scores judges from first-token logits, measured position bias is inflated by about 42 points. Generate the verdict before auditing a judge.