Models#Benchmark#Model inference
Claude Haiku 5.5: cheapest, fastest yet
$0.10/$0.50 per million tokens, about 75% cheaper than 4.5, with OSWorld at 72.4% and a first effort dial.

Anthropic released Claude Haiku 5.5 (claude-haiku-5-5) on October 7, positioned as “the cheapest, fastest, and most capable small model we’ve ever released” — aimed at high-volume, cost-sensitive work (summaries, compaction, database queries, classification) and at serving as the subagent alongside Opus 5.5 and Sonnet 5.5. It is live now on AWS, Google Cloud, Azure and the Claude platform.
Key points
- Pricing: $0.10 input / $0.50 output per million tokens at the ≤100k tier (Haiku 4.5 was $1.00/$5.00); average running cost about 75% below 4.5, 90% below for the requests that make up 90% of traffic
- Benchmarks: OSWorld 2.1 offline at 72.4% (4.5: 15.7%); Terminal-Bench 4.0 at 39.2% (4.5: 0%); Humanity’s Last Exam 45.9% without tools
- New: the first Haiku with an adjustable effort dial (Low–Max); SDK betas for computer and browser use
- Bundled: Sonnet 5.5 cache reads halved to $0.10; monthly API credits for Max and Team subscribers ($100–$500)
- Customers: Asana reports 30%+ latency reduction and up to 2.5x faster inference per agent turn; Cognition’s Devin Fusion scored 66.2 on FrontierCode with Haiku 5.5 as its sidekick
- Timing: multiple outlets note the launch lands ahead of a reported November IPO at a $2 trillion valuation
The arithmetic of the price war
The $0.10/$0.50 tag sits directly against GPT-6 Luna — VentureBeat’s headline was “90% price reduction, matching GPT-6 Luna.” Small models are the consumable of the agent economy: subagents invoke them every turn, so unit price decides product margins. Cutting the inference floor by three quarters on the eve of an IPO with a leaked $8B operating loss is a unit-economics answer delivered to the public market before the prospectus.
A generational jump for small models
Going from 15.7% to 72.4% on OSWorld in one generation says the small-model ceiling is rising at unprecedented speed: last generation’s Haiku summarized and classified, this one drives computers and terminals. The safety posture moved too — cyber safeguards are tighter than 4.5 (penetration testing is blocked outright) while biology safeguards match the frontier models, a fine-grained split that mirrors where small models actually get misused. The effort dial concedes a related reality — one model must now trade cost per token against task depth on demand, which is exactly what agent schedulers need.
Customer numbers and real workloads
Anthropic unusually led with customer measurements: Asana reports latency down more than 30% with up to 2.5x faster inference per agent turn; HubSpot scored 92.8 on its CRM suite; Box gained 11 points of quality at roughly half the latency; and Cognition’s Devin Fusion hit 66.2 on FrontierCode with Haiku 5.5 as its sidekick. The common thread is the subagent pattern — a frontier model plans, Haiku runs the errands, and the bill is dominated by how many errands there are. That is precisely why the small-model price war matters: the marginal cost of the agent economy is this price times call volume.
One caveat
The benchmarks are vendor-chosen tracks; third-party reproduction will take weeks. And the “75% cheaper” average leans on the claim that 90% of requests sit under 100k tokens — long-context agents that read whole codebases fall into the 50% discount tier above it. Cost-sensitive teams should map their own traffic distribution against the two tiers before concluding anything from a leaderboard.
A three-way repricing, in one week
Haiku 5.5 ($0.10/$0.50), GPT-6 Luna, and Mistral Large 4 ($1.36/$4.18) staged a full price-ladder recalibration inside seven days. For developers, subagent cost models now have a shelf life measured in quarters; for the industry, a pre-IPO Anthropic has pinned its unit-economics story to the most aggressive price cut of the season. The counter-reading is just as plausible: with Cognition’s FrontierCode 66.2 as proof, cheap-and-good subagents may widen the moat rather than compress it — every agent built on Haiku 5.5 is an agent that will not be rebuilt on a rival’s stack at the same price.