~/satyajit

Claude Haiku 5.5: what "75% less to run" is made of

mdjsonmcp

2026-10-08 · 28 min · small-models · pricing · benchmarks · inference · tokenization · computer-use

Why read this

Notabletop 60%

Splits Haiku 5.5's '75% cheaper' into price, tokenizer, tier and effort, with per-task costs decoded from Anthropic's own charts and a list-price calculator.

  • A guide you can follow today
  • Something most practitioners use
  • Original analysis

Inference & servingAPI onlyProprietaryPractitioner model

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
1 of 3: API-only, gated or restrictive licence
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
3 of 3: A decision guide a practitioner can follow today
Will it last?
1 of 3: Relevant for months
Does it affect many?
3 of 3: Something most practitioners touch
Only here?
2 of 3: A teardown or measurement few others did

Score 66 of 100, ranked 163 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Anthropic released Claude Haiku 5.5 on October 7, and the post that announced it led with one number: "On average, it costs around 75% less to run than Claude Haiku 4.5." A cost claim with "on average" in it is a claim about a mix of workloads, and I wanted to know which mix, because the answer for a classifier and the answer for an agent loop are not going to be the same number.

A disclosure before anything else: this site is written and maintained with Claude agents, so Anthropic is my vendor. I have tried to read their numbers the way I would read anyone's, which mostly means going to the footnotes.

The short version is that "75% less" is the product of three things that pull in different directions, plus a fourth that the headline benchmarks quietly depend on. The per-token price fell by 90% or 50% depending on prompt length. The tokenizer changed, so the same text is about 30% more tokens. The model now thinks by default, which is more output tokens. And the benchmark table at the top of the launch was run at max effort, where Haiku 5.5 spends so many tokens that on three of four of Anthropic's own cost charts it costs more per task than Haiku 4.5 did. At the default effort it is far cheaper on all four. Both of those are true, and which one applies to you is a setting you choose.

The claim, and the footnote under it

The announcement's footnote 2 says what the 75% is built from. I am quoting it in full, because every word in it carries part of the calculation:

Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category. This calculation also accounts for changes between Haiku 4.5 and Haiku 5.5 in how many tokens are used to complete a given piece of work: Haiku 5.5 has an updated tokenizer (similar to Sonnet 5.5's and Opus 5.5's), which means it uses slightly more tokens per task.

So the 75% is a per-task cost, averaged over the traffic Anthropic saw on Haiku 4.5, with a token-count adjustment folded in. Three inputs: the price cut in each tier, the share of traffic in each tier, and how many more tokens Haiku 5.5 uses for the same job.

Pricing table, price per one million tokens. Haiku 5.5, prompts up to or over 100k: cache reads 0.01 or 0.05 dollars, cache writes 0.125 or 0.625, input 0.10 or 0.50, output 0.50 or 2.50. Haiku 4.5: cache reads 0.10, cache writes 1.25, input 1.00, output 5.00. Sonnet 5.5: cache reads 0.10, cache writes 2.50, input 2.00, output 10.00.
The launch price table: Haiku 5.5 has two columns, split at a 100,000-token prompt. Every Haiku 5.5 price is exactly a tenth of Haiku 4.5's in the first column and exactly half in the second (Anthropic, Claude Haiku 5.5 announcement).

The price ratio is the easy part

I checked the table against the pricing page in Anthropic's docs, which has two things the launch graphic leaves out: the 1-hour cache write ($0.20 and $1 per million tokens for the two tiers, against $2 on Haiku 4.5) and the Batch API ($0.05 and $0.25 input, $0.25 and $1.25 output, against $0.50 and $2.50). Every one of them follows the same rule. In the low tier, each Haiku 5.5 price is 0.1 times the Haiku 4.5 price. In the high tier, each is 0.5 times. Input, output, cache reads, 5-minute and 1-hour cache writes, batch: no exceptions.

That uniformity is the most useful fact on the page, because it collapses the cost comparison into one line. If rpr_p is the price ratio for the tier a request lands in and rtr_t is how many tokens Haiku 5.5 uses for the job relative to Haiku 4.5, then

C5.5C4.5=rp⋅rt,rp∈{0.1, 0.5}\frac{C_{5.5}}{C_{4.5}} = r_p \cdot r_t, \qquad r_p \in \{0.1,\ 0.5\}

as long as the mix of token types stays the same. The whole argument about whether Haiku 5.5 is cheaper for you comes down to rtr_t, and rtr_t is where everything interesting happens.

The tokenizer, which is not "slightly"

The footnote calls the token increase slight. The model docs are more specific. The Haiku 5.5 overview page, the "What's new" page and the migration guide all say the same sentence: Haiku 5.5 "uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5." The general pricing page says the same of every model on that tokenizer, and adds that the exact increase depends on the content.

Thirty percent is not slight. On its own it turns the low tier's 90% cut into 87% (0.1×1.3=0.130.1 \times 1.3 = 0.13) and the high tier's 50% cut into 35% (0.5×1.3=0.650.5 \times 1.3 = 0.65). I suspect the footnote's "slightly more tokens per task" is a net figure, with the tokenizer's 30% partly offset by a model that does the same job in fewer words or fewer turns, but the footnote does not say that and I could not check it.

The tokenizer also moves the tier boundary. The 100,000-token threshold is counted in Haiku 5.5's tokens, so a prompt that Haiku 4.5 counted as about 77,000 tokens (100,000/1.3100{,}000 / 1.3) is already over the line on Haiku 5.5. If you sized a pipeline's chunks to stay well under 100k on the old model, recount them. The migration guide says as much: "Count your prompts with model set to claude-haiku-5-5 rather than reusing counts measured on Claude Haiku 4.5."

The 100k cliff

The high tier is not marginal pricing. Once a prompt is over 100,000 tokens, every token in that request, the first 100,000 included, is billed at the higher rate. The pricing page puts it as "a prompt of over 100,000 tokens pays higher prices," and the system card describes its own cost accounting the same way: "Claude Haiku 5.5 is charged per request at the rate for that request's prompt length (above or below 100K tokens)." So the cost of a request jumps fivefold at the boundary. A 99,000-token prompt with a 1,000-token answer costs about $0.0104. Add 2,000 tokens to the prompt and it costs about $0.0530.

This is new for Anthropic's current line. The same pricing page says every Claude 4.6-or-later model "except Claude Haiku 5.5" bills the full 1M context at standard rates. The low tier is aimed squarely at the competition: Artificial Analysis noted that $0.10 and $0.50 is the same as GPT-6 Luna's price.

One thing I could not pin down: whether cached tokens count toward the 100,000. The docs say "prompt" and never define it for this purpose. My calculator below counts input, cache reads and cache writes together, which is the conservative reading. If you run a long cached system prompt on Haiku 5.5, measure your own bill before trusting either of us.

Rebuilding the 75%

With the price ratios and the tokenizer factor known, I tried to get from the footnote's inputs back to its output. The footnote says 90% of Haiku 4.5 requests were under 100k. If you weight by requests, ignoring that long requests cost more, the per-task ratio with a 1.3x tokenizer is

0.9×0.13+0.1×0.65=0.1820.9 \times 0.13 + 0.1 \times 0.65 = 0.182

which is 82% less, not 75%. Weighting by requests is the wrong weighting, though. A 150,000-token request costs ten times what a 15,000-token one does, so the 10% of long requests carry far more than 10% of the spend. If ss is the share of Haiku 4.5 spend that sat in requests over 100k, the blended ratio is 0.13(1−s)+0.65s0.13(1-s) + 0.65s, and it equals 0.25 when ss is about 0.23. With no tokenizer penalty at all (rt=1r_t = 1), it takes ss of about 0.375.

So the 75% is consistent with somewhere between a quarter and two-fifths of Haiku 4.5's spend having been in long requests, depending on how much the token count really grew. That is a plausible traffic mix, and nothing in it looks wrong. It is also a number about Anthropic's aggregate traffic. Your number depends on where your own spend sits, and the calculator is the fastest way to find it.

Your workload, priced from the list

The calculator below uses nothing but the published per-million-token prices. You enter a request the way you already measure it, in Haiku 4.5 tokens; it scales the counts by the tokenizer factor for Haiku 5.5 and Sonnet 5.5, adds thinking tokens on the newer models, picks Haiku 5.5's tier from the scaled prompt length, and prices all three.

cost per request · published list pricestokens as Haiku 4.5 counts them

A cached system prompt and history, a short reply, a little thinking.

Haiku 5.5 vs Haiku 4.5: 84.3% cheaper

Every Haiku 5.5 price is a tenth of Haiku 4.5's in the low tier and half of it in the high tier, across input, output, cache reads and cache writes alike. So the ratio between the two models is just the price ratio times the token ratio. The tokenizer pushes the token ratio up by about 1.3x for the same text, thinking pushes it up further, and a prompt that crosses 100,000 Haiku 5.5 tokens moves the whole request to the 5x tier.

The presets are workloads I picked to cover the shapes people actually run. With the 1.3x tokenizer and the list prices, per request:

Workload (Haiku 4.5 tokens)Haiku 4.5Haiku 5.5changeSonnet 5.5
Ticket triage: 1,200 in, 15 out$0.001275$0.000166-87%$0.003315
Support chat: 2,000 in, 18,000 cache read, 350 out, 300 thinking$0.00555$0.000871-84%$0.01509
Summarise 70,000 in, 1,200 out$0.076$0.00988-87%$0.1976
Summarise 85,000 in, 1,200 out$0.091$0.05915-35%$0.2366
Subagent: 3,000 in, 45,000 cache read, 3,000 cache write, 1,500 out, 2,500 thinking$0.01875$0.003688-80%$0.0679

The two summarisation rows are the cliff in one picture. The documents differ by 15,000 tokens. On Haiku 4.5 that is a 20% difference in cost. On Haiku 5.5 the 85,000-token document becomes 110,500 tokens, crosses the line, and costs six times what the 70,000-token one does. If your documents sit in that band, chunking them below the threshold is worth more than any prompt trimming you will ever do.

Thinking barely registers in the support-chat and subagent rows, because output in the low tier is $0.50 per million. It matters much more at high effort, which is where the benchmarks were run.

Which effort the benchmarks used

Haiku 5.5 is the first Haiku with an effort setting (low, medium, high, xhigh, max) and adaptive thinking, which is on by default. The default effort on the API is medium. Here is the benchmark table from the launch.

Benchmark table comparing Haiku 5.5, Haiku 4.5, GPT-6 Luna and Sonnet 5.5 for reference. GDPval-AA v2.1: 1620, 735, 1437, 1840. AA-Briefcase v1.1: 1578, 614, 1336, 1824. OSWorld 2.1 offline subset: 72.4%, 15.7%, 48.9%, 83.9%. Humanity's Last Exam no tools: 45.9%, 10.2%, dash, 56.9%. With tools: 57.4%, 18.7%, dash, 64.5%. Terminal-Bench 4.0: 39.2%, 0.0%, 16.4%, 70.6%. FrontierCode 1.1 Main: 46.4%, dash, 42.4%, 52.1% at xhigh. Chartography no tools: 46.4%, 6.4%, 29.1%, 61.6%.
The launch benchmark table. The system card's version of it (Table 8.1.A) states the configuration: every Haiku 5.5 number is at max effort, averaged over five trials (Anthropic, Claude Haiku 5.5 announcement).

The launch page does not say which effort produced these numbers. The system card does, under its Table 8.1.A: "all Haiku 5.5 results use the following standard configuration: adaptive thinking at max effort." That is a normal way to report a model's ceiling. It matters here because the cost headline and the capability headline then describe different settings, and the announcement shows how far apart those settings are in three charts of score against cost per task.

Scatter chart, GDPval-AA v2.1 Elo against cost per task in USD on a log scale. Haiku 5.5 rises from Low near 1.2 cents and 1125 Elo through Med, High and Xhigh to Max near 87 cents and 1620 Elo. Sonnet 5.5 runs from about 22 cents to about 6.8 dollars. GPT-6 Luna runs from under half a cent to about 9 cents. Haiku 4.5 is a single point near 24 cents and 735 Elo.
GDPval-AA by effort level. Haiku 4.5's single point sits at about 24 cents a task; Haiku 5.5's max point, the one in the table, sits to its right (Anthropic, Claude Haiku 5.5 announcement).

Look at where the Haiku 4.5 point sits on the cost axis, then at Haiku 5.5's max point. The max point is to the right of it. Max-effort Haiku 5.5 costs more per GDPval task than Haiku 4.5 did.

I wanted the numbers rather than my eye on a log axis. The charts on Anthropic's page are inline SVG: every point is a <circle> with pixel coordinates and every axis label is positioned text, so fitting the tick positions (log for cost, linear for score) turns each circle back into a cost and a score. The decode is good to about a pixel. It reads the GDPval medium point as 1275 Elo, and the system card gives 1277 for that setting; it reads OSWorld max as 72.4% and Terminal-Bench max as 39.2%, which match the table exactly. The widget holds what came out.

cost per task by effort · Haiku 5.5 vs Haiku 4.5decoded from Anthropic's charts
$0.01$0.1$1cost per task, USD, log scaleHaiku 4.5 · $0.240 · 735 Elolow · $0.012 · -95% · 1125 Elomedium · $0.030 · -87% · 1275 Elohigh · $0.089 · -63% · 1418 Eloxhigh · $0.268 · +12% · 1512 Elomax · $0.867 · +261% · 1620 Elo

Green points cost less per task than Haiku 4.5, orange points cost more. Each label gives the effort level, the cost per task, the change against Haiku 4.5, and the score. The launch table quotes the max-effort score. On three of these four benchmarks, max effort costs more per task than Haiku 4.5 did. The default effort is medium.

Here are the default and the max points against Haiku 4.5, cost per task with the score beside it:

BenchmarkHaiku 4.5Haiku 5.5, mediumHaiku 5.5, max
OSWorld 2.1 (offline subset)$1.45 · 15.7%$0.126 · 53.3% (-91%)$0.61 · 72.4% (-58%)
GDPval-AA v2.1$0.24 · 735 Elo$0.031 · 1275 (-87%)$0.87 · 1620 (3.6x)
Humanity's Last Exam, no tools$0.094 · 10.2%$0.0069 · 35.5% (-93%)$0.21 · 45.9% (2.3x)
Terminal-Bench 4.0$0.79 · 0.0%$0.68 · 20.2% (-14%)$2.65 · 39.2% (3.4x)

OSWorld is the one benchmark where every effort level beats Haiku 4.5 on both axes. The Terminal-Bench numbers come from a fourth chart, in the announcement's section on Sonnet 5.5's price cut, which plots Haiku 4.5 and 5.5 alongside it.

Scatter chart, OSWorld 2.1 offline subset partial-credit score against cost per attempt in USD on a log scale. Haiku 5.5 rises from Low at about 7 cents and 42% to Max at about 61 cents and 72%. GPT-6 Luna runs from about 4 cents and 19% to about 21 cents and 49%. Sonnet 5.5 runs from about 68 cents and 58% to about 5.8 dollars and 84%. Haiku 4.5 sits alone near 1.45 dollars and 16%.
OSWorld 2.1 by effort level: here every Haiku 5.5 setting is both cheaper and better than Haiku 4.5's single point (Anthropic, Claude Haiku 5.5 announcement).

Two caveats run in opposite directions, and I think both are fair. First, Haiku 4.5 has no effort setting, so the system card ran it with a large fixed thinking budget: 63,999 tokens on Terminal-Bench, 60,000 on the HLE-with-tools run. Most production Haiku 4.5 traffic runs with thinking off, so this baseline is Haiku 4.5 at its most expensive, and a like-for-like saving against your real Haiku 4.5 bill could be smaller than these charts show. Second, Haiku 4.5 scores 0.0% on Terminal-Bench and 10.2% on HLE at that cost. A cheaper model that fails is not a bargain, and the comparison that matters on those rows is "medium Haiku 5.5 does a job Haiku 4.5 could not."

So the headline is not wrong, and it is not the whole story either. At medium, on these four charts, Haiku 5.5 is 14% to 93% cheaper per task than Haiku 4.5 and far better. At max, the setting behind every number in the table, it costs more on three of the four. "75% cheaper" and "scores 1620 on GDPval" are each true, and they are not true at the same time.

Scatter chart, Humanity's Last Exam without tools, score against cost per attempt on a log scale. Haiku 5.5 runs from Low near a third of a cent and 31% to Max near 21 cents and 46%. Sonnet 5.5 runs from about 3 cents and 41% to about 90 cents and 57%. Haiku 4.5 sits alone near 9 cents and 10%.
Humanity's Last Exam without tools by effort. The step from xhigh to max adds under two points and roughly triples the cost (Anthropic, Claude Haiku 5.5 announcement).

The HLE chart also shows how flat the top of the curve is: xhigh at 44.1% for about $0.065, max at 45.9% for about $0.21. On the healthcare evaluations in the system card, response time on HealthBench Professional was "8 to 21 seconds per answer from low to xhigh and about 110 seconds at max." Max effort is a setting for a benchmark table or for the occasional request that really needs it, and the docs say so: for xhigh and max, "also run your evals on Claude Sonnet 5.5 and compare performance, cost, and speed."

Where the tokens go

The per-task costs above are the price ratio times the token ratio, so a cost ratio above 1 at a price ratio of 0.1 means Haiku 5.5 used more than ten times the tokens. Agent loops make that plausible, because their prompts grow turn after turn and a long run drifts over 100,000 tokens, where the price ratio is 0.5 instead. I cannot split GDPval's 3.6x into tier and tokens from the chart alone; it implies max-effort Haiku 5.5 spent somewhere between about 7 and 36 times Haiku 4.5's tokens, depending on how much of the run billed in the high tier.

Two independent pictures of token use agree on the shape. Artificial Analysis counts output tokens per task across its Intelligence Index.

Bar chart of output tokens per Artificial Analysis Intelligence Index task. Claude 4.5 Haiku uses 18k. Claude Haiku 5.5 uses 17k at low, 33k at medium, 55k at high, 89k at xhigh and 162k at max. GPT-6 Luna uses 2k at low, 11k at medium, 20k at high, 27k at xhigh and 50k at max.
Output tokens per Intelligence Index task: Haiku 4.5 at 18k, Haiku 5.5 from 17k at low to 162k at max (Artificial Analysis, launch-day thread).

Haiku 4.5 used about 18,000 output tokens per task. Haiku 5.5 uses about 17,000 at low, 33,000 at medium and 162,000 at max. On output alone, in the low tier, medium costs 33,000 × $0.50 per million against 18,000 × $5, about 82% less; max costs 162,000 × $0.50, about 10% less, and if those prompts were over 100k it is 162,000 × $2.50, several times more. Artificial Analysis also put Haiku 5.5 at max at "~3x GPT-6 Luna (max, ~50k)" and noted that at a similar score of 38, Haiku 5.5 at high used about 55,000 tokens to Luna's 50,000. Its cost-per-task figures were provisional on launch day, because its site did not yet model the tiered pricing.

The system card's FrontierCode chart puts output tokens on the x-axis too, and it shows the same knee.

Line chart, FrontierCode Main score against average output tokens per task on a log scale. Claude Haiku 5.5 rises from low near 24k tokens and 35% to medium near 36k and 42%, high near 55k and 42%, xhigh near 100k and 46%, and max near 180k and 46%. Claude Sonnet 5.5 peaks near 52% at about 48k tokens and falls to 46% at about 580k at max. Other lines show GPT-6 Sol, GPT-6 Luna, Claude Opus 5, Claude Opus 5.5 and Claude Fable 5.1.
FrontierCode Main against output tokens per task, run by Cognition. Haiku 5.5's max point, 46.4%, uses nearly twice the tokens of its xhigh point at 45.8% (Claude Haiku 5.5 system card, Figure 8.3.A).

FrontierCode is a nice illustration of a different point. The table compares Haiku 5.5's 46.4% at max with Sonnet 5.5's 52.1% at xhigh, Sonnet's best setting. At max Sonnet 5.5 scored 46.2%, basically tied with Haiku 5.5, after spending something like 580,000 output tokens per task. More effort is not monotonically better on every benchmark, for either model.

Benchmarks, and who ran them

The table compares Haiku 5.5 with Haiku 4.5, GPT-6 Luna and, "for reference," Sonnet 5.5. The system card's summary adds SWE-bench (Pro 64.8%, Multilingual 83.7% against Haiku 4.5's 67.4%, Multimodal 30.7% against 19.8%) and HealthBench Professional (64.8% against 32.2%). Who ran each number matters more than usual, because several are third-party:

Artificial Analysis's launch-day summary is the most useful outside view. It puts Haiku 5.5 at 43 on its Intelligence Index at max effort, up from 17 for Haiku 4.5, ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38), and 13 points behind Sonnet 5.5 (56). It also flagged a weak spot: AA-Omniscience knowledge accuracy of 36%, against 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna, with a lower hallucination rate (40% against 55% and 77%). A small model that knows less and says so more often is a reasonable trade for the jobs Haiku is sold for.

Two results surprised me. On ProgramBench, a long-context coding benchmark whose episodes run up to the full 1M window, Haiku 5.5 scored 82.0% against Sonnet 5.5's 79.7%. And Anthropic is unusually blunt about the opposite case: "Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0," where Sonnet 5.5 scored 70.6%.

Speed: "fastest" without a tokens-per-second number

The announcement calls Haiku 5.5 "our fastest model to date," and the footnote narrows it to "at each model's standard speed, although it runs less quickly than our Opus models in Fast Mode." Anthropic publishes no tokens-per-second figure; the models overview just lists its comparative latency as "Fastest."

The measured numbers come from Artificial Analysis, a day after launch: 137 output tokens per second at medium, 160 at high, 178 at low, 187 at xhigh and 242 at max, counted after the first chunk arrives. Its pages show about 90 for Haiku 4.5 and 101 to 129 for Sonnet 5.5. Two adjustments before reading that as "50% faster." The tokens are new-tokenizer tokens, so 137 of them carry roughly as much text as 105 of Haiku 4.5's. And with thinking on, the time you wait is mostly thinking: Artificial Analysis lists 9.93 seconds to the first answer token at low effort. Why max effort streams fastest I don't know; it is their measurement, one day old.

The customer quotes are all relative and unverifiable from outside. Asana reported "over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn" against "the model we use today," which it does not name. Box reported Haiku 5.5 "scored 11 points higher than Haiku 4.5 at about half the latency." If you need latency, measure time to first visible token at the effort you will actually run; the effort setting moves it more than the model choice does.

Context, limits and the migration

The spec sheet, from the Haiku 5.5 model page in the docs:

Haiku 4.5Haiku 5.5
Model IDclaude-haiku-4-5-20251001claude-haiku-5-5
Context window200k1M
Max output64k128k (300k on the Batch API with a beta header)
Thinkingmanual budgetadaptive, on by default
Default effortnonemedium
Reliable knowledge cutoffFebruary 2025June 2026
Retirementnot sooner than October 15, 2026not sooner than October 7, 2027

The retirement row is the one to read twice. Haiku 4.5's commitment runs only to October 15, 2026, a week after this launch, so the migration is not optional for long.

There is a system card, 144 pages, and the capability numbers above come from its section 8. The migration guide lists the breaking changes, and three of them touch cost or correctness directly. Thinking counts toward max_tokens, so a small max_tokens sized for Haiku 4.5 can stop after a thinking block with no text. temperature, top_p and top_k must be left at defaults or the request returns a 400. Assistant prefill returns a 400 too, so classifiers that prefilled a label need structured outputs or a tool with an enum instead. Responses can also end with stop_reason: "refusal" from safety classifiers, with no server-side fallback, and Priority Tier is not supported.

The prompting guide has one more cost note worth knowing: in long agent prompts, "moving from low to medium effort roughly halved early stopping. It also more than doubled the output tokens for each attempt."

The other price cut

The thread ended with "one more thing": Sonnet 5.5's cache reads drop from $0.20 to $0.10 per million, which Anthropic says makes it "around 20% cheaper to run on most long-running work." The Terminal-Bench chart plots Sonnet 5.5 at both prices, so I decoded that too. Per attempt, the new price is 21% to 26% cheaper across the five effort levels, about $14.11 to $10.45 at max. For that drop, cache reads must have been roughly 40% to 50% of Sonnet 5.5's spend on those runs, which is about what you would expect from an agent rereading its own transcript every turn. The claim holds on Anthropic's own chart.

27.8 billion tokens for $20?

A day after launch, @1kartikkabadi1 posted that "you can get ~27.8b tokens on the $20 plan on Haiku 5.5" and linked a page that works it out. My first reaction was that the arithmetic cannot come from the API. Twenty dollars at Haiku 5.5's cheapest price, the $0.01 cache read, buys 2 billion tokens. Getting to 27.8 billion needs something other than the price list, so I read the page.

It is careful, and it shows its work. It treats a Pro subscription as a fixed monthly pool of usage valued at API list prices, and divides that pool by a blended price per million tokens. The pool is $1,280, which comes from a public project that reads usage panels from saturated Pro accounts: $3.20 of list-price Opus 5.5 usage per percentage point of the weekly meter, times 100 points, times four weeks. The blend assumes 97% of tokens are cache reads, 2.5% are 5-minute cache writes and 0.5% are output, a mix the page takes from measured agentic-coding sessions. On Haiku 5.5's short-prompt prices that is 0.97×0.01+0.025×0.125+0.005×0.50=0.0153250.97 \times 0.01 + 0.025 \times 0.125 + 0.005 \times 0.50 = 0.015325 dollars per million, and $1,280 divided by it is 83.5 billion tokens. At the long-prompt prices the blend is five times higher and the answer is 16.7 billion. The 27.8 is neither: it is a 50/50 mix, $1,280 over the average of the two blends. Every price the page uses matches the table above.

Dot chart titled Tokens per month on Claude Pro, 1,280 dollar anchor, log scale from 1B to 100B tokens. Opus 5.5 at 3.05B, Sonnet 5.5 pre-cut at 4.17B, Sonnet 5.5 post-cut at 6.11B, Haiku 5.5 over 100k at 16.7B, Haiku 5.5 50/50 mix at 27.8B, Haiku 5.5 up to 100k at 83.5B.
The page's estimate of what a Pro plan buys at saturated use, one row per model and tier. The 27.8B in the post is the middle Haiku row, a 50/50 split across the 100k threshold (claude-55-lineup.vercel.app, 'The 5.5 lineup on $20 per month').

So the post's number is $20 multiplied up 64 times by a pool estimate, then divided by a price that is 97% cache reads. Both of those deserve a sentence.

The pool was measured on Opus 5.5. Applying it to Haiku assumes the plan's meter charges each model in proportion to its API list price, so that a percentage point is $3.20 of Haiku just as it is $3.20 of Opus. Anthropic publishes no token allowances, and nobody had a saturated Haiku 5.5 week to measure on the day it shipped. If the meter weighs Haiku more heavily than its price, every Haiku row shrinks, and I have no way to tell from outside. The page itself also leads with the lower of two pool measurements; the other, from Opus 5 accounts in September, is 55% higher. One reply to the post reports 2.2 to 2.5 billion Opus 5.5 tokens a week on a single account, against the page's 0.76 billion. That is one person, but it says the spread is wide.

The cache-read share decides what a "token" is here. Of the 27.8 billion, 27.0 billion are cache reads, the same context re-sent on every turn of an agent loop and billed at a tenth of the input price. Only 139 million are output. That is a fair description of how Claude Code burns tokens, and a misleading one if you read 27.8 billion as text the model writes or reads fresh. A classifier with short, uncached prompts has a blend nearer the $0.10 input price, which turns the same pool into about 13 billion tokens, and nobody runs a classifier through a chat subscription anyway.

Then there is the cliff from earlier. For a workload that is 97% cache reads, the question I could not settle, whether cached tokens count toward the 100,000, is the whole answer: if they do, any agent whose context grows past 100k pays the long-prompt rate on every turn, and the right row is 16.7, not 83.5. The poster's replies say he runs Claude Code with autocompaction at 100k, which is exactly the habit that keeps a session in the cheap tier under the conservative reading. And the 50/50 row splits tokens, not money: at five times the price, the long half eats five-sixths of the pool.

What survives is the ratio. With one pool for every model, Haiku 5.5 at short prompts comes out about 27 times Opus 5.5, because its blended price is about a twenty-seventh of Opus's. That is a real statement about relative cost on the subscription, if the meter follows list prices, and it is the same statement the price table already makes. The absolute 27.8 billion is a ceiling stacked on two estimates, the pool and the traffic mix, plus an unpublished assumption about how Haiku is metered. I would plan around the ratio and treat the headline as an upper bound I have not seen anyone hit.

What I'd do with it

For classification, extraction, routing and short summaries, the migration is close to free money: about 87% off at the same text, as long as you keep prompts under the threshold, pick low or medium effort, and turn thinking off where the task does not need it (it can be disabled at high effort and below). For anything with long prompts, recount in Haiku 5.5 tokens first. Documents that land between roughly 77,000 and 100,000 Haiku 4.5 tokens are where the tier boundary bites, and chunking them is the biggest single saving available. For agent loops, the price is not the variable; effort is. Medium is where the "75% less" lives. Max is where the benchmark table lives, and on most of Anthropic's own charts it costs more than the model it replaces.

The related pieces on this site go at the same question from other sides: the harness effect on why orchestration sets an agent's token bill more than the model does, how LLM inference works for why prefill and decode cost what they do, and tiny browser models, which priced a Haiku 4.5 call against a 28 KB model; at Haiku 5.5's prices that gap shrinks about tenfold.

How I checked

I read the announcement page, its footnotes and the X thread from @claudeai through the fxtwitter mirror, the 144-page system card (sections 1, 2 and 8 in full), and the docs pages for the models overview, Haiku 5.5 overview, what's new, migration guide, pricing, context windows, effort, prompt caching and the Haiku 5.5 prompting guide, all on October 8. The per-task costs come from decoding the four inline-SVG charts on the announcement page: I took each axis's tick positions, fitted a log scale for cost and a linear one for score, and mapped every <circle> back to data. I checked the decode against the system card's own numbers (GDPval medium 1275 against 1277, OSWorld and Terminal-Bench max exact). Workload costs are list prices times the token counts stated, with a 1.3x tokenizer factor from the docs; that factor is Anthropic's approximation, and the real one depends on your text. Speed figures are Artificial Analysis's, read from its model pages, and I did not measure latency myself. I did not call the API for this piece, so the threshold question, whether cached tokens count toward 100,000, is unverified. For the subscription estimate I read the post and its replies through the fxtwitter mirror and rendered the lineup page in a headless browser; its blends and token counts recompute exactly from its stated formula and prices, and its $1,280 pool is the page's own sourced estimate, which I could not check.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Claude Haiku 5.5: what "75% less to run" is made of", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026claudehaiku55,
  author = {Satyajit Ghana},
  title  = {Claude Haiku 5.5: what "75% less to run" is made of},
  url    = {https://ai.thesatyajit.com/articles/claude-haiku-5-5},
  year   = {2026}
}
share