Vibeleaderboard
← All Intel
Intel / post

Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953…

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys

Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B (Reasoning), 1141 for Gemini 3.5 Flash-Lite, and 1004 for Kimi K2.5 (Reasoning), with Terminal-Bench v2.1 (52%) also behind Qwen3.6 27B (61%). The hallucination gap follows: 82% on AA-Omniscience against 49% for Qwen3.6 27B and 34% for Flash-Lite (lower is better). Its size-twin Gemma 4 31B (Reasoning) performs worse than Muse Glimmer on all of these measures. The exception is agentic tool use, with Muse Glimmer scoring strongly on Tau3-Banking (24%), ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%)

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
  • agentic loop — The cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.
Why it matters

The headline intelligence score hides a large agentic and gap — a reason not to swap this model into an on the index number alone.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…