
The 3-point gain over Muse Spark 1.1 on the Artificial Analysis Intelligence Index is concentrated in agentic evaluations: GDPval-AA v2 +260 Elo (1371 to 1631), Terminal-Bench 2.1 +2 points (78% to 80%), and Tau3-Bench Banking +2 points (25% to 27%). The minor regressions are SciCode (-2 points) and Humanity's Last Exam (-1 point)

Seeing the gain concentrated in agentic evaluations while other capabilities flatten tells you what the model is actually better at, which is what matters when slotting it into an .
Checking sign-in…
Loading comments…