Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
Source
Haowei Liu, Hsin-Tai Wu, Yi Fang
Author
Haowei Liu, Hsin-Tai Wu, Yi Fang
Date
Key takeaways · AI-distilled
Measured against two-author gold labels, a deployed gpt-4o-mini text-to-SQL judge reached Cohen's kappa of only 0.04 on a disagreement-enriched set and 0.42 on a random spot-check, over-flagging 77.1% of faithful cases.
A self-hosted Qwen3.6-27B judge reached kappa 0.72, close to Claude Opus 4.7 at 0.71, at roughly 1/300 of the cost per call; the authors note the head-to-head is underpowered at n=96.
Pairing a weak judge with a strong one lowered agreement, while three strong judges with unanimity routing reached kappa 0.79 and auto-handled 89.7% of cases.
Applied to BIRD-financial, the same audit recipe flagged 25.5% of the expert-written gold SQL queries as candidate gold-label issues under the authors' protocol.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
Why it matters
A deployed GPT-4o-mini judge in a production text-to-SQL pipeline agreed with human graders at kappa 0.04, over-flagging 77% of correct results; swapping to a self-hosted 27B model matched Claude Opus 4.7 at roughly 1/300th the cost.