AA-Omniscience regresses 10 points from Qwen3.7 Max (+14 to +4), reversing its predecessor's abstention gains. Accuracy is effectively flat at ~31% while the hallucination rate rises from 23% to 40%, meaning Qwen3.8 Max attempts more questions it cannot answer rather than declining them
hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
agentic loop — The cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.
Why it matters
Anyone considering Qwen3.8 Max in an agentic loopThe cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.Full definition → needs to know its hallucinationWhen a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.Full definition → rate nearly doubled while accuracy stayed flat. A model that stopped abstaining is materially riskier in unattended pipelines than its headline index score suggests.