The 3-point gain over Muse Spark 1.1 on the Artificial Analysis Intelligence Index is concentrated in agentic evaluations: GDPval-AA v2 +260 Elo (1371 to 1631), Terminal-Bench 2.1 +2 points (78% to 80%), and Tau3-Bench Banking +2 points (25% to 27%). The minor regressions are SciCode (-2 points) and Humanity's Last Exam (-1 point)
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
Why it matters
Seeing the gain concentrated in agentic evaluations while other capabilities flatten tells you what the model is actually better at, which is what matters when slotting it into an agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition →.