
Muse Spark 1.2 continues Meta's AA-Omniscience pattern: the score rose from 18 to 22, driven by abstention for the second consecutive release. The hallucination rate fell 10 points (38% to 28%) as the attempt rate dropped to 67%, while accuracy slipped from 41% to 38%. Heavy abstention drives both the low hallucination rate and the lower accuracy

Knowing that a model's honesty score improved by answering less often rather than knowing more changes how you'd deploy it in retrieval or loops where abstention costs you a step.
Checking sign-in…
Loading comments…