Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 El
Source
ArtificialAnlys
Author
ArtificialAnlys
Published
Terms in this piece · Glossary
hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
agentic loop — The cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.
Why it matters
The headline intelligence score hides a large agentic and hallucinationWhen a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.Full definition → gap — a reason not to swap this model into an agentic loopThe cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.Full definition → on the index number alone.
Transcript
Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B (Reasoning), 1141 for Gemini 3.5 Flash-Lite, and 1004 for Kimi K2.5 (Reasoning), with Terminal-Bench v2.1 (52%) also behind Qwen3.6 27B (61%). The hallucination gap follows: 82% on AA-Omniscience against 49% for Qwen3.6 27B and 34% for Flash-Lite (lower is better). Its size-twin Gemma 4 31B (Reasoning) performs worse than Muse Glimmer on all of these measures. The exception is agentic tool use, with Muse Glimmer scoring strongly on Tau3-Banking (24%), ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%)