Anthropic releases Claude Haiku 5.5: large benchmark gains over Haiku 4.5 and roughly 75% lower average running cost
Synopsis
Anthropic released Claude Haiku 5.5, a small model for high-volume, cost-sensitive tasks, reporting gains over Haiku 4.5 across GDPval-AA v2.1, AA-Briefcase v1.1, OSWorld 2.1, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode 1.1 and Chartography, a first-for-Haiku adjustable effort setting, and roughly 75% lower average running cost.
Interpretation
Haiku 5.5 scores well above Haiku 4.5 on multiple benchmarks: knowledge work GDPval-AA v2.1 at 1620 versus 735 and AA-Briefcase v1.1 at 1578 versus 614; computer use OSWorld 2.1 offline subset at 72.4% versus 15.7%; multidisciplinary reasoning Humanity's Last Exam at 45.9% without tools and 57.4% with tools versus 10.2% and 18.7%; agentic coding Terminal-Bench 4.0 at 39.2% versus 0.0% and FrontierCode 1.1 (Main) at 46.4%; visual reasoning Chartography at 46.4% versus 6.4%. Relative to the previous Haiku generation, the reported gains span knowledge work, computer use, reasoning, coding and visual reasoning rather than a single task; the text also lists GPT-6 Luna and Sonnet 5.5 for reference, with Sonnet 5.5 still higher on most entries (for example 70.6% on Terminal-Bench 4.0). Evidence is the benchmark table in the text, with evaluation details deferred to the Haiku 5.5 System Card; OSWorld 2.1 and Chartography are labeled offline subset or no tools, Humanity's Last Exam distinguishes with and without tools, and FrontierCode 1.1 marks Sonnet 5.5 as Xhigh.
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, letting users choose between optimizing for cost or intelligence; the text says charts show its performance on three benchmarks at each effort setting. This setting was not previously available in the Haiku class and had been offered on other models, so a single model can now be tuned between cost and capability per task. The text presents the three-benchmark results per effort setting as charts without listing the values in prose, and points to the System Card for detail.
On pricing, Haiku 5.5 costs $0.01 per million tokens for cache reads on prompts up to 100k and $0.05 over 100k, $0.125 / $0.625 for cache writes, $0.10 / $0.50 for input tokens and $0.50 / $2.50 for output tokens; Haiku 4.5 is listed at $0.10, $1.25, $1.00 and $5.00, and Sonnet 5.5 at $0.10, $2.50, $2.00 and $10.00. The text states Haiku 5.5 costs around 75% less to run on average than Haiku 4.5; a footnote specifies 90% lower for requests up to 100,000 tokens and 50% lower above that, notes that 90% of Haiku 4.5 requests fell in the former category, and says the calculation accounts for the updated tokenizer using slightly more tokens per task. Evidence is the pricing table and footnote in the text comparing list prices and describing the request distribution, as published by the vendor.
On safety, the text reports major improvements across almost all alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior and lower willingness to cooperate with misuse; cybersecurity safeguards are more restrictive than Haiku 4.5's but somewhat less restrictive than other recent models, permitting a wider range of defensive tasks than Sonnet 5.5 while still blocking penetration testing and techniques more likely used by attackers; biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5, allowing research biology questions while restricting requests judged likely to cause harm. Alignment evaluations improve overall while safeguard strength is differentiated by capability, with application routes through the Life Sciences Verification Program and Cyber Verification Program for wider-ranging biology and cyber activities. Evidence is the qualitative description of alignment evaluations and safeguard policy in the text, with the evaluation process and results deferred to the System Card.
Perspective
The release targets developers and teams needing high-volume, cost-sensitive inference: summaries, compactions, database queries, classification requests, and use as a subagent alongside Opus 5.5 and Sonnet 5.5 on coding work; it also suits speed-sensitive settings such as live customer support and browser use. The text states Haiku 5.5 is best suited to more narrowly scoped tasks that might otherwise have been cost-prohibitive, while Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0. Availability covers AWS, Google Cloud, Microsoft Azure and the Claude Platform, with developers starting via claude-haiku-5-5 and a migration guide. Accompanying updates include halving Sonnet 5.5 cache read prices from $0.20 to $0.10 per million tokens, reducing cost on most agentic tasks by around 20%, and a monthly API credit for Max and Team subscribers (Max 5x at $100 per month, Max 20x at $200, Team up to $500 pooled across users).
Benchmark numbers are vendor-reported; the text gives no sample sizes, evaluation protocols or confidence intervals, and defers evaluation detail to the System Card. OSWorld 2.1 and Chartography are labeled offline subset or no tools, Humanity's Last Exam distinguishes with and without tools, and FrontierCode 1.1 marks Sonnet 5.5 as Xhigh, so these condition differences matter when comparing across models. Per-effort-setting values on three benchmarks are shown only as charts, so the size of the cost-versus-capability tradeoff at each level cannot be judged from the text. Customer feedback is summarized only as consistent with the reported performance and cost improvements, without specific data. The safety section is qualitative, and specific alignment metrics and results require the System Card.
