
If replicated, this means math-reasoning pipelines can swap expensive frontier judges for a cheap three-model unanimous-vote ensemble at up to 100x lower cost without losing reliability.
“On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agree with human pass/fail decisions at rates statistically indistinguishable from Claude Opus 4.7 and Gemini 3.1 Pro, at up to $100\times$ lower cost.”
“We had expected a majority vote of the three to be the best budget option; it matched the frontier but did not improve on its strongest member.”
“The headline finding is that cheap judges are competitive with the frontier at one to two orders of magnitude lower cost; as a deployable default we recommend all-three-pass, with the caveat that this rule was identified post-hoc and warrants independent replication.”
articleAre the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon StatementsXinke Tong, Xuanming Zhang, Tianyi Tang, An Yang, Jiatu Hu, Guojie Lin, Zhenzhen Shi, Lingfeng Zeng, Boyu Yang, Bing Zhao, Hu Wei, Lin Qu, Dayiheng Liu
articleWhen LLM judges agree, should we believe them?www.amazon.scienceChecking sign-in…
Loading comments…