Anthropic's latest entry in the high-speed, low-cost tier, Claude Haiku 5.5, has been added to Artificial Analysis with a dedicated profile detailing the metrics used to evaluate it. The platform has released a description of the methodology behind its v4.3.2 Intelligence Index, outlining how it measures performance across agentic workflows, cost efficiency, and latency. This update clarifies the specific benchmarks and calculation methods applied to the model, providing transparency on how Haiku 5.5's capabilities are quantified.

A New Standard for Intelligence Measurement

The core of this analysis relies on the Artificial Analysis Intelligence Index v4.3.2, which aggregates ten distinct evaluations. These include AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The source text explicitly lists these components, noting that AA-Briefcase v1.1 is an agentic knowledge work benchmark. Its associated metric, AA-Briefcase Elo, is a combined score that aggregates rubric pass rate, analytical quality Elo, and presentation Elo. Additionally, the index incorporates AA-Omniscience, which measures knowledge reliability and hallucination by rewarding correct answers and penalizing hallucinations, with scores ranging from -100 to 100. Terminal-Bench 4.0 is specifically utilized for agentic scientific research workflows in a terminal.

Cost Efficiency and Token Economics

For builders operating at scale, the cost per Intelligence Index task is a primary metric defined by the platform. Artificial Analysis calculates this weighted average cost by determining the cost for each evaluation based on input, cache hit, cache write, reasoning, and answer token prices, dividing by task count, and weighting by each benchmark's importance in the Intelligence Index. The methodology page also defines 'Cost to Run Artificial Analysis Intelligence Index' as the total USD cost to run all evaluations, calculated using the model's specific token prices and usage. While these definitions establish how cost is measured, the source text provided does not contain the actual comparative data or competitor rankings for Haiku 5.5.

Speed, Latency, and End-to-End Response

The profile details the methodology for measuring speed and latency. Output Speed is defined as tokens per second received while the model is generating tokens, representing performance on the first-party API or the median across providers. Time per Intelligence Index Task is calculated as the weighted average decode time per task, excluding time to first token and overhead, derived by dividing output tokens per task by output speed. Furthermore, End-to-End Response Time is defined as the seconds required to output 500 tokens, which includes time to first token, 'thinking' time for reasoning models, and answer generation time. The source text presents these as definitions of the metrics available for analysis, rather than results indicating a specific strategic focus on minimizing decode time.

Key Takeaways

  • The AA Intelligence Index v4.3.2 includes ten specific evaluations, including AA-Briefcase v1.1 and GDPval-AA v2.1, with the source text explicitly listing version numbers for these benchmarks.
  • AA-Briefcase Elo aggregates rubric pass rate, analytical quality, and presentation, while AA-Omniscience measures knowledge reliability by penalizing hallucinations.
  • The source text defines the methodology for calculating weighted average cost and end-to-end response time but does not provide the actual comparative performance data or competitor rankings.

The Bottom Line

This update is a transparency report, not a results analysis. It clarifies exactly how Artificial Analysis measures Claude Haiku 5.5's intelligence, cost, and speed, but developers must wait for the actual data points to see how the model stacks up against competitors.