Vibeleaderboard
← All Intel
Intel / post

Hallucination gating collapses legal agent benchmark scores

Source
x.com
Date
ArtificialAnlys@ArtificialAnlys

Today we are announcing Harvey LAB-AA v1.1 in collaboration with Harvey. This updates our scoring methodology for the Legal Agent Benchmark (LAB) to add a hallucination check and require correct responses to not include material misstatements. LAB-AA v1.1's new headline metric, Hallucination-Gated All-Pass Rate, only credits a task when the deliverables satisfy every rubric criterion and contain no material hallucinations. We define a material hallucination as one that would mislead a reader on a substantive point, such as a wrong contractually required date, while a minor hallucination is a real error that is unlikely to meaningfully affect the legal interpretation of a deliverable. Minor hallucinations are reported separately and do not affect the headline score. Grok 4.7 (xhigh) leads at 9.4% Hallucination-Gated All-Pass Rate, and >60% of otherwise passing results across the models tested at launch contain a material hallucination. This is the first step in enhancing the methodology for Harvey LAB-AA. In future updates, we’re working with Harvey to better account for the full set of factors lawyers value, including usability features like style and tone. Key takeaways: ➤…

Read the full post on X
Why it matters

Gating scores on material hallucinations cuts the top legal-agent score to 9.4%, and over 60% of otherwise passing outputs contain one. Pass rates alone overstate reliability.

Key takeaways · AI-distilled
  • LAB-AA v1.1 credits a legal task only when deliverables meet every rubric criterion and contain no material , defined as an error that would mislead on a substantive point, such as a wrong contractually required date. Minor errors are reported separately.
  • The gate reorders the board: Muse Spark 1.3 (max) would lead at a 26.7% all-pass rate without it, but two thirds of those passes contain a material hallucination, leaving 8.9%. GPT-6 Astra (max) keeps almost all its passes (8.9% to 8.6%) and rises from joint 10th to 3rd.
  • Per Artificial Analysis, GPT-6 Astra (max) averaged 0.03 material hallucinations per task (4 across 120 tasks) and had none on a 20-task subset under all six checker models, while Gemini 3.8 Flash (high) averaged 13.96 per task.
  • Top score did not track price: Grok 4.7 (xhigh) leads at about $9.50 per task, under half the roughly $21.70 of Claude Fable 5.1 (max with fallback), the most expensive model tested, while Muse Spark 1.3 (max) placed second at about $4.20 per task.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…