
“On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agree with human pass/fail decisions at rates statistically indistinguishable from Claude Opus 4.7 and Gemini 3.1 Pro, at up to $100\times$ lower cost.”
“We had expected a majority vote of the three to be the best budget option; it matched the frontier but did not improve on its strongest member.”
“The headline finding is that cheap judges are competitive with the frontier at one to two orders of magnitude lower cost; as a deployable default we recommend all-three-pass, with the caveat that this rule was identified post-hoc and warrants independent replication.”
articleChain-of-Models: Cross-Model Auditing for Bias-Robust LLM JudgesQian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Bingsheng He
articleEvaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)eugeneyan.comSign in to comment.
Loading comments…