
Qwen3.8 Max scores 1739 Elo on GDPval-AA, ahead of Kimi K3 (1685), effectively tied with Claude Fable 5 (1743) and GPT-5.6 Sol (max, 1730), and behind only Claude Opus 5 (max, 1852). The 468 Elo gain is in part driven by the new model taking more turns per task (64 vs. 14 for Qwen3.7 Max)

A gain driven by taking 4x more turns per task is a different purchase than a gain in per-turn capability: it lands as latency and spend in production. This is the kind of read a headline alone would hide.
Checking sign-in…
Loading comments…