Anthropic launches Claude Haiku 5.5: OSWorld jumps from 15.7% to 72.4%, pricing drops up to 90%

Official Anthropic announcement image for Claude Haiku 5.5.Anthropic
Official Anthropic announcement image for Claude Haiku 5.5.Anthropic
AI & Automation

Claude Haiku 5.5 is live on the Claude Platform, AWS, Google Cloud and Azure. The new Haiku runs about 75% cheaper than Haiku 4.5, scores 72.4% on OSWorld-2.1, gains adjustable effort settings, and arrives with a Sonnet 5.5 cache-read cut and new monthly API credits.

Anthropic released Claude Haiku 5.5 on October 7, 2026, calling it its fastest, cheapest and most capable small model. The gains are not incremental: on the OSWorld-2.1 computer-use benchmark (offline subset), Haiku 5.5 scores 72.4% against Haiku 4.5's 15.7%, and on Terminal-Bench 4.0 for agentic coding it reaches 39.2% where its predecessor scored 0%. Average running cost drops about 75% compared with Haiku 4.5.

The model is available immediately under the ID claude-haiku-5-5 on the Claude Platform, and through Amazon Web Services, Google Cloud and Microsoft Azure. The launch follows Claude Sonnet 5.5 from September 28 and rounds out Anthropic's 5.5 lineup below Opus 5.5.

A cheap subagent that stopped being bad at computer use

Anthropic positions Haiku 5.5 for high-volume, cost-sensitive work: summaries, compactions, database queries, classification, live customer support and browser use. The official announcement also recommends it as a subagent paired with Opus 5.5 or Sonnet 5.5 on coding work.

The benchmark table in the announcement puts Haiku 5.5 at 1,620 on GDPval-AA v2.1 versus 735 for Haiku 4.5, and 45.9% on Humanity's Last Exam without tools (up from 10.2%). Against OpenAI's budget model GPT-6 Luna, Anthropic lists Haiku 5.5 leading every tested category. The same table shows Sonnet 5.5 at 83.9% on OSWorld and 70.6% on Terminal-Bench — this is a strong small model, not a mid-tier replacement. The Decoder independently corroborates the launch, the benchmark jumps and the pricing structure.

The pricing cut has a tokenizer caveat

Published per-million-token prices: input $0.10 and output $0.50 for prompts up to 100,000 tokens; input $0.50 and output $2.50 above that threshold. Cache reads cost $0.01 (up to 100k) or $0.05, versus $1.00 input and $5.00 output on Haiku 4.5. Anthropic says short-prompt traffic makes up roughly 90% of previous Haiku requests, so effective savings there reach up to 90%.

Two caveats are worth keeping. First, The Decoder notes Haiku 5.5 ships with an updated tokenizer that consumes slightly more tokens per task than Haiku 4.5 — per-token price cuts therefore overstate real-world savings. Second, prompts over 100,000 tokens cost five times as much as short ones. Long-context jobs priced on Haiku 4.5 numbers need re-measuring before anyone banks the headline discount.

First Haiku with adjustable effort, plus credits for Max and Team

Haiku 5.5 is the first Haiku-class model with an adjustable effort setting (Low through Max), letting developers trade cost against intelligence instead of accepting one fixed operating point. Anthropic publishes accuracy-versus-cost curves for OSWorld, GDPval-AA and Humanity's Last Exam, and the Haiku 5.5 System Card documents the evaluation details.

Two platform changes ship with it. Sonnet 5.5 cache reads are halved to $0.10 per million tokens, cutting most agentic workloads by roughly 20% because cache reads dominate token consumption. New monthly API credits roll out for Claude Max and Team subscribers; The Decoder reports $100 for Max 5x, $200 for Max 20x and up to $500 for Team, intended for building agents and applications on the Claude Platform. Check the Claude pricing page for the current tier terms.

Early customer numbers, and what safeguards changed

Anthropic quotes enterprise testers: Asana reports over 30% lower task-completion latency and up to 2.5x faster inference per agent turn; HubSpot recorded its best score yet on a CRM task suite at 92.8% averaged over three runs; Box measured 11 points above Haiku 4.5 at about half the latency; AlphaSense saw a statistically significant quality gain on 400 sampled production queries. These are vendor-relayed customer claims, not independent benchmarks.

Safeguards shifted with capability. Cybersecurity restrictions are tighter than Haiku 4.5's but looser than Sonnet 5.5's: a wider range of defensive tasks is permitted, while penetration testing stays blocked. Biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5. Broader research access runs through Anthropic's Life Sciences and Cyber Verification Programs.

Migrating without relearning your costs

Anthropic links an official migration guide and a support article covering availability. The practical migration risk is not model quality but cost modelling: the new tokenizer, the 100k prompt threshold and per-workload effort settings all change the arithmetic of a Haiku fleet. Run your own eval set at the effort level you plan to ship, measure tokens per task against Haiku 4.5, and then decide whether the price cut holds for your mix.

The pressure works both ways — The Decoder frames the Sonnet 5.5 cache cut as a response to OpenAI's GPT-6.1 series, and LinkLoot has tracked OpenAI's parallel moves (GPT-6.1 Sol). For teams running summarization, classification or browser-agent pipelines at volume, the open question now is which small model still fits the budget after tokenizer effects. More agent-building workflows are collected in our AI agent tools guide.

From reading to doing

Try the related loot

Best provider for OpenClaw in 2026: what to buy, what to avoid, and what actually saves money

Open loot