
Seeing the gain concentrated in agentic evaluations while other capabilities flatten tells you what the model is actually better at, which is what matters when slotting it into an .
postQwen3.8 Max costs $1.14 per Intelligence Index task, more than double Qwen3.7 Ma
postQwen3.8 Max scores 1739 Elo on GDPval-AA, ahead of Kimi K3 (1685), effectively t
postAlibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at
postAA-Omniscience regresses 10 points from Qwen3.7 Max (+14 to +4), reversing its pSign in to comment.
Loading comments…