Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953…
- Source
- Artificial Analysis
- Date

Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B (Reasoning), 1141 for Gemini 3.5 Flash-Lite, and 1004 for Kimi K2.5 (Reasoning), with Terminal-Bench v2.1 (52%) also behind Qwen3.6 27B (61%). The hallucination gap follows: 82% on AA-Omniscience against 49% for Qwen3.6 27B and 34% for Flash-Lite (lower is better). Its size-twin Gemma 4 31B (Reasoning) performs worse than Muse Glimmer on all of these measures. The exception is agentic tool use, with Muse Glimmer scoring strongly on Tau3-Banking (24%), ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%)

- eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
- hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
- agentic loop — The cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.
The headline intelligence score hides a large agentic and gap — a reason not to swap this model into an on the index number alone.
Checking sign-in…
Loading comments…





