Vibeleaderboard
← All Intel
Intel / post

Artificial Analysis: Haiku 5.5 leads small models but burns tokens

Source
x.com
Date
ArtificialAnlys@ArtificialAnlys

Anthropic has released Claude Haiku 5.5, scoring 43 on the Artificial Analysis Intelligence Index - up 26 points one year after the last Haiku release Haiku 5.5 is the first Haiku model with Anthropic’s effort settings and adaptive thinking, and Anthropic has introduced tiered pricing. Haiku 5.5 is cheaper than its predecessor - it costs $0.10/$0.50 per 1M input/output tokens for prompts up to 100k tokens (the same as GPT-6 Luna and 10% of the previous Haiku model). However, this pricing rises 5x to $0.50/$2.50 above 100k. The site does not yet reflect tiered pricing, so provisional cost figures for Haiku 5.5 do not include the step up cost. We are working on support and will follow up with Cost per Task coverage soon. Key takeaways: ➤ Leading small-class model performance: At max effort Haiku 5.5 sits slightly ahead of models such as GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38). Its score is comparable to Kimi K3 (44), a 2.8T parameter open weights model, and trails Claude Sonnet 5.5 (max, 56) by 13 points ➤ Heavy token use compared to GPT-6 Luna: Haiku 5.5 (max) uses ~162k output tokens per Intelligence Index task, ~3x GPT-6 Luna (max, ~50k). Moving from…

Read the full post on X
Why it matters

Haiku 5.5 scores 43 on the Intelligence Index, but its price rises 5x above 100k tokens and it uses about 3x the output tokens of GPT-6 Luna at max effort. Real cost per task may erase the headline price advantage.

Key takeaways · AI-distilled
  • At max effort Haiku 5.5 scores 43 on the Artificial Analysis Intelligence Index, ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38), close to Kimi K3 (44) and 13 points behind Claude Sonnet 5.5 at max (56).
  • Moving Haiku 5.5 from xhigh to max adds 2 index points for about 1.8x the . At matched scores of 38, Haiku 5.5 at high uses about 55k tokens per task against about 50k for Luna at max, and AA says the gap widens at lower effort.
  • On Terminal-Bench 4.0 Haiku 5.5 scores 33%, up from 0% for Haiku 4.5 and ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%). On AA-Briefcase knowledge work it reaches 1578 Elo, ahead of Kimi K3 and GLM-5.3.
  • Its AA-Omniscience accuracy is 36% versus 55% for Gemini 3.8 Flash and 44% for Luna, partly because it more often admits not knowing: its rate is 40% versus 55% and 77%.
  • AA calls its 35% AutomationBench-AA score likely understated because a pre-release safety refusal issue caused over-refusal, and will re-run after Anthropic's fix. Its provisional cost figures also omit the above-100k price step.
Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…