
Benchmarks 13 current models as PR reviewers on real bugs and shows GPT-5.2 and Opus 4.6 lead on quality, but cheaper models like Kimi K2.5 hit 75-86% of that quality for a fraction of the per-PR cost, a real model-selection tradeoff.
Checking sign-in…
Loading comments…