
AA-Omniscience regresses 10 points from Qwen3.7 Max (+14 to +4), reversing its predecessor's abstention gains. Accuracy is effectively flat at ~31% while the hallucination rate rises from 23% to 40%, meaning Qwen3.8 Max attempts more questions it cannot answer rather than declining them

Anyone considering Qwen3.8 Max in an needs to know its rate nearly doubled while accuracy stayed flat. A model that stopped abstaining is materially riskier in unattended pipelines than its headline index score suggests.
postMuse Spark 1.2 continues Meta's AA-Omniscience pattern: the score rose from 18…Artificial Analysis
postSolar Pro 4's AA-Omniscience score improves from -53 to -1, however the improvemArtificialAnlys
postLing 3.0 Flash scores -18 on AA-Omniscience, a 48 point improvement from Ling…Artificial AnalysisChecking sign-in…
Loading comments…