Artificial Analysis scores Claude Haiku 5.5 at 43 on its Intelligence Index
Artificial Analysis scored Claude Haiku 5.5 at 43 on its Intelligence Index, 26 points above the previous Haiku one year earlier. The context window is 1 million tokens, up from 200,000 for Claude 4.5 Haiku. The model takes text and image input and produces text output.
At maximum effort, the score is slightly ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38), comparable to Kimi K3 (44), and 13 points behind Claude Sonnet 5.5 at maximum effort (56). On the same index, maximum effort uses about 162,000 output tokens per task, roughly three times GPT-6 Luna at maximum effort (about 50,000). Moving from xhigh to max adds 2 points for about 1.8 times the tokens. At a matching score of 38, high effort uses about 55,000 tokens per task, against about 50,000 for GPT-6 Luna at maximum effort.
On AA-Briefcase, Artificial Analysis’s private evaluation of realistic knowledge-work tasks, maximum effort reaches 1,578 Elo, ahead of Kimi K3 and GLM-5.3 and comparable to Muse Spark 1.3 at maximum effort. On Terminal-Bench 4.0 it scores 33%, up from 0% for Haiku 4.5, level with GLM-5.3 Flash and ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%).
AA-Omniscience accuracy is 36%, compared with 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna. Artificial Analysis says part of the gap comes from a greater willingness to admit when it does not know: the hallucination rate is 40%, against 55% and 77%. On AutomationBench-AA the score is 35%, against 53–60% for those three models. Artificial Analysis says a safety-refusal issue in pre-release testing caused over-refusal, Anthropic is working on a fix, and the firm will re-run the evaluation and expects the score to rise.
Five-minute cache writes are $0.125 per million tokens for prompts up to 100,000 tokens and $0.625 above that. Artificial Analysis says its site does not yet reflect tiered pricing, so provisional cost figures for Haiku 5.5 omit the higher rate above 100,000 tokens.