
It places a new frontier-scale model on the cost/capability frontier against the open-weights leader, which is the comparison an engineer makes when choosing a model to build on.
postQwen3.8 Max costs $1.14 per Intelligence Index task, more than double Qwen3.7 Ma
postQwen3.8 Max scores 1739 Elo on GDPval-AA, ahead of Kimi K3 (1685), effectively t
postAA-Omniscience regresses 10 points from Qwen3.7 Max (+14 to +4), reversing its p
postOnly two labs occupy the Time per Task Pareto frontier: all frontier models undeSign in to comment.
Loading comments…