Muse Spark 1.2 ranks #5 on GDPval-AA v2, our benchmark of agentic real-world knowledge work, at 1631 Elo - behind e.g. Claude Opus 5 (max, 1852), GPT-5.6 Sol (max, 1730), and Kimi K3 (1685), and ahead of e.g. Claude Opus 4.8 (max, 1588). Muse Spark 1.1 scored 1371 at its launch last month
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters
Direct Elo comparison against the models practitioners already run is the fastest way to judge whether a new coding AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → is worth trialing.