Vibeleaderboard
← All Intel
Intel / post

Muse Spark 1.2 ranks #5 on GDPval-AA v2, our benchmark of agentic real-world…

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys

Muse Spark 1.2 ranks #5 on GDPval-AA v2, our benchmark of agentic real-world knowledge work, at 1631 Elo - behind e.g. Claude Opus 5 (max, 1852), GPT-5.6 Sol (max, 1730), and Kimi K3 (1685), and ahead of e.g. Claude Opus 4.8 (max, 1588). Muse Spark 1.1 scored 1371 at its launch last month

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Direct Elo comparison against the models practitioners already run is the fastest way to judge whether a new coding is worth trialing.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…