2026-10-08 · 28 min · small-models · pricing · benchmarks · inference · tokenization · computer-use
Why read this
Notabletop 60%Splits Haiku 5.5's '75% cheaper' into price, tokenizer, tier and effort, with per-task costs decoded from Anthropic's own charts and a list-price calculator.
- A guide you can follow today
- Something most practitioners use
- Original analysis
Inference & servingAPI onlyProprietaryPractitioner model
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 1 of 3: API-only, gated or restrictive licence
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 3 of 3: A decision guide a practitioner can follow today
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 3 of 3: Something most practitioners touch
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 66 of 100, ranked 163 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
Anthropic released Claude Haiku 5.5 on October 7, and the post that announced it led with one number: "On average, it costs around 75% less to run than Claude Haiku 4.5." A cost claim with "on average" in it is a claim about a mix of workloads, and I wanted to know which mix, because the answer for a classifier and the answer for an agent loop are not going to be the same number.
A disclosure before anything else: this site is written and maintained with Claude agents, so Anthropic is my vendor. I have tried to read their numbers the way I would read anyone's, which mostly means going to the footnotes.
The short version is that "75% less" is the product of three things that pull in different directions, plus a fourth that the headline benchmarks quietly depend on. The per-token price fell by 90% or 50% depending on prompt length. The tokenizer changed, so the same text is about 30% more tokens. The model now thinks by default, which is more output tokens. And the benchmark table at the top of the launch was run at max effort, where Haiku 5.5 spends so many tokens that on three of four of Anthropic's own cost charts it costs more per task than Haiku 4.5 did. At the default effort it is far cheaper on all four. Both of those are true, and which one applies to you is a setting you choose.
The claim, and the footnote under it
The announcement's footnote 2 says what the 75% is built from. I am quoting it in full, because every word in it carries part of the calculation:
Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category. This calculation also accounts for changes between Haiku 4.5 and Haiku 5.5 in how many tokens are used to complete a given piece of work: Haiku 5.5 has an updated tokenizer (similar to Sonnet 5.5's and Opus 5.5's), which means it uses slightly more tokens per task.
So the 75% is a per-task cost, averaged over the traffic Anthropic saw on Haiku 4.5, with a token-count adjustment folded in. Three inputs: the price cut in each tier, the share of traffic in each tier, and how many more tokens Haiku 5.5 uses for the same job.

The price ratio is the easy part
I checked the table against the pricing page in Anthropic's docs, which has two things the launch graphic leaves out: the 1-hour cache write ($0.20 and $1 per million tokens for the two tiers, against $2 on Haiku 4.5) and the Batch API ($0.05 and $0.25 input, $0.25 and $1.25 output, against $0.50 and $2.50). Every one of them follows the same rule. In the low tier, each Haiku 5.5 price is 0.1 times the Haiku 4.5 price. In the high tier, each is 0.5 times. Input, output, cache reads, 5-minute and 1-hour cache writes, batch: no exceptions.
That uniformity is the most useful fact on the page, because it collapses the cost comparison into one line. If is the price ratio for the tier a request lands in and is how many tokens Haiku 5.5 uses for the job relative to Haiku 4.5, then
as long as the mix of token types stays the same. The whole argument about whether Haiku 5.5 is cheaper for you comes down to , and is where everything interesting happens.
The tokenizer, which is not "slightly"
The footnote calls the token increase slight. The model docs are more specific. The Haiku 5.5 overview page, the "What's new" page and the migration guide all say the same sentence: Haiku 5.5 "uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5." The general pricing page says the same of every model on that tokenizer, and adds that the exact increase depends on the content.
Thirty percent is not slight. On its own it turns the low tier's 90% cut into 87% () and the high tier's 50% cut into 35% (). I suspect the footnote's "slightly more tokens per task" is a net figure, with the tokenizer's 30% partly offset by a model that does the same job in fewer words or fewer turns, but the footnote does not say that and I could not check it.
The tokenizer also moves the tier boundary. The 100,000-token threshold is counted
in Haiku 5.5's tokens, so a prompt that Haiku 4.5 counted as about 77,000 tokens
() is already over the line on Haiku 5.5. If you sized a pipeline's
chunks to stay well under 100k on the old model, recount them. The migration guide
says as much: "Count your prompts with model set to claude-haiku-5-5 rather than
reusing counts measured on Claude Haiku 4.5."
The 100k cliff
The high tier is not marginal pricing. Once a prompt is over 100,000 tokens, every token in that request, the first 100,000 included, is billed at the higher rate. The pricing page puts it as "a prompt of over 100,000 tokens pays higher prices," and the system card describes its own cost accounting the same way: "Claude Haiku 5.5 is charged per request at the rate for that request's prompt length (above or below 100K tokens)." So the cost of a request jumps fivefold at the boundary. A 99,000-token prompt with a 1,000-token answer costs about $0.0104. Add 2,000 tokens to the prompt and it costs about $0.0530.
This is new for Anthropic's current line. The same pricing page says every Claude 4.6-or-later model "except Claude Haiku 5.5" bills the full 1M context at standard rates. The low tier is aimed squarely at the competition: Artificial Analysis noted that $0.10 and $0.50 is the same as GPT-6 Luna's price.
One thing I could not pin down: whether cached tokens count toward the 100,000. The docs say "prompt" and never define it for this purpose. My calculator below counts input, cache reads and cache writes together, which is the conservative reading. If you run a long cached system prompt on Haiku 5.5, measure your own bill before trusting either of us.
Rebuilding the 75%
With the price ratios and the tokenizer factor known, I tried to get from the footnote's inputs back to its output. The footnote says 90% of Haiku 4.5 requests were under 100k. If you weight by requests, ignoring that long requests cost more, the per-task ratio with a 1.3x tokenizer is
which is 82% less, not 75%. Weighting by requests is the wrong weighting, though. A 150,000-token request costs ten times what a 15,000-token one does, so the 10% of long requests carry far more than 10% of the spend. If is the share of Haiku 4.5 spend that sat in requests over 100k, the blended ratio is , and it equals 0.25 when is about 0.23. With no tokenizer penalty at all (), it takes of about 0.375.
So the 75% is consistent with somewhere between a quarter and two-fifths of Haiku 4.5's spend having been in long requests, depending on how much the token count really grew. That is a plausible traffic mix, and nothing in it looks wrong. It is also a number about Anthropic's aggregate traffic. Your number depends on where your own spend sits, and the calculator is the fastest way to find it.
Your workload, priced from the list
The calculator below uses nothing but the published per-million-token prices. You enter a request the way you already measure it, in Haiku 4.5 tokens; it scales the counts by the tokenizer factor for Haiku 5.5 and Sonnet 5.5, adds thinking tokens on the newer models, picks Haiku 5.5's tier from the scaled prompt length, and prices all three.
A cached system prompt and history, a short reply, a little thinking.
Every Haiku 5.5 price is a tenth of Haiku 4.5's in the low tier and half of it in the high tier, across input, output, cache reads and cache writes alike. So the ratio between the two models is just the price ratio times the token ratio. The tokenizer pushes the token ratio up by about 1.3x for the same text, thinking pushes it up further, and a prompt that crosses 100,000 Haiku 5.5 tokens moves the whole request to the 5x tier.
The presets are workloads I picked to cover the shapes people actually run. With the 1.3x tokenizer and the list prices, per request:
| Workload (Haiku 4.5 tokens) | Haiku 4.5 | Haiku 5.5 | change | Sonnet 5.5 |
|---|---|---|---|---|
| Ticket triage: 1,200 in, 15 out | $0.001275 | $0.000166 | -87% | $0.003315 |
| Support chat: 2,000 in, 18,000 cache read, 350 out, 300 thinking | $0.00555 | $0.000871 | -84% | $0.01509 |
| Summarise 70,000 in, 1,200 out | $0.076 | $0.00988 | -87% | $0.1976 |
| Summarise 85,000 in, 1,200 out | $0.091 | $0.05915 | -35% | $0.2366 |
| Subagent: 3,000 in, 45,000 cache read, 3,000 cache write, 1,500 out, 2,500 thinking | $0.01875 | $0.003688 | -80% | $0.0679 |
The two summarisation rows are the cliff in one picture. The documents differ by 15,000 tokens. On Haiku 4.5 that is a 20% difference in cost. On Haiku 5.5 the 85,000-token document becomes 110,500 tokens, crosses the line, and costs six times what the 70,000-token one does. If your documents sit in that band, chunking them below the threshold is worth more than any prompt trimming you will ever do.
Thinking barely registers in the support-chat and subagent rows, because output in the low tier is $0.50 per million. It matters much more at high effort, which is where the benchmarks were run.
Which effort the benchmarks used
Haiku 5.5 is the first Haiku with an effort setting (low, medium, high, xhigh, max) and adaptive thinking, which is on by default. The default effort on the API is medium. Here is the benchmark table from the launch.

The launch page does not say which effort produced these numbers. The system card does, under its Table 8.1.A: "all Haiku 5.5 results use the following standard configuration: adaptive thinking at max effort." That is a normal way to report a model's ceiling. It matters here because the cost headline and the capability headline then describe different settings, and the announcement shows how far apart those settings are in three charts of score against cost per task.

Look at where the Haiku 4.5 point sits on the cost axis, then at Haiku 5.5's max point. The max point is to the right of it. Max-effort Haiku 5.5 costs more per GDPval task than Haiku 4.5 did.
I wanted the numbers rather than my eye on a log axis. The charts on Anthropic's
page are inline SVG: every point is a <circle> with pixel coordinates and every
axis label is positioned text, so fitting the tick positions (log for cost, linear
for score) turns each circle back into a cost and a score. The decode is good to
about a pixel. It reads the GDPval medium point as 1275 Elo, and the system card
gives 1277 for that setting; it reads OSWorld max as 72.4% and Terminal-Bench max as
39.2%, which match the table exactly. The widget holds what came out.
Green points cost less per task than Haiku 4.5, orange points cost more. Each label gives the effort level, the cost per task, the change against Haiku 4.5, and the score. The launch table quotes the max-effort score. On three of these four benchmarks, max effort costs more per task than Haiku 4.5 did. The default effort is medium.
Here are the default and the max points against Haiku 4.5, cost per task with the score beside it:
| Benchmark | Haiku 4.5 | Haiku 5.5, medium | Haiku 5.5, max |
|---|---|---|---|
| OSWorld 2.1 (offline subset) | $1.45 · 15.7% | $0.126 · 53.3% (-91%) | $0.61 · 72.4% (-58%) |
| GDPval-AA v2.1 | $0.24 · 735 Elo | $0.031 · 1275 (-87%) | $0.87 · 1620 (3.6x) |
| Humanity's Last Exam, no tools | $0.094 · 10.2% | $0.0069 · 35.5% (-93%) | $0.21 · 45.9% (2.3x) |
| Terminal-Bench 4.0 | $0.79 · 0.0% | $0.68 · 20.2% (-14%) | $2.65 · 39.2% (3.4x) |
OSWorld is the one benchmark where every effort level beats Haiku 4.5 on both axes. The Terminal-Bench numbers come from a fourth chart, in the announcement's section on Sonnet 5.5's price cut, which plots Haiku 4.5 and 5.5 alongside it.

Two caveats run in opposite directions, and I think both are fair. First, Haiku 4.5 has no effort setting, so the system card ran it with a large fixed thinking budget: 63,999 tokens on Terminal-Bench, 60,000 on the HLE-with-tools run. Most production Haiku 4.5 traffic runs with thinking off, so this baseline is Haiku 4.5 at its most expensive, and a like-for-like saving against your real Haiku 4.5 bill could be smaller than these charts show. Second, Haiku 4.5 scores 0.0% on Terminal-Bench and 10.2% on HLE at that cost. A cheaper model that fails is not a bargain, and the comparison that matters on those rows is "medium Haiku 5.5 does a job Haiku 4.5 could not."
So the headline is not wrong, and it is not the whole story either. At medium, on these four charts, Haiku 5.5 is 14% to 93% cheaper per task than Haiku 4.5 and far better. At max, the setting behind every number in the table, it costs more on three of the four. "75% cheaper" and "scores 1620 on GDPval" are each true, and they are not true at the same time.

The HLE chart also shows how flat the top of the curve is: xhigh at 44.1% for about $0.065, max at 45.9% for about $0.21. On the healthcare evaluations in the system card, response time on HealthBench Professional was "8 to 21 seconds per answer from low to xhigh and about 110 seconds at max." Max effort is a setting for a benchmark table or for the occasional request that really needs it, and the docs say so: for xhigh and max, "also run your evals on Claude Sonnet 5.5 and compare performance, cost, and speed."
Where the tokens go
The per-task costs above are the price ratio times the token ratio, so a cost ratio above 1 at a price ratio of 0.1 means Haiku 5.5 used more than ten times the tokens. Agent loops make that plausible, because their prompts grow turn after turn and a long run drifts over 100,000 tokens, where the price ratio is 0.5 instead. I cannot split GDPval's 3.6x into tier and tokens from the chart alone; it implies max-effort Haiku 5.5 spent somewhere between about 7 and 36 times Haiku 4.5's tokens, depending on how much of the run billed in the high tier.
Two independent pictures of token use agree on the shape. Artificial Analysis counts output tokens per task across its Intelligence Index.

Haiku 4.5 used about 18,000 output tokens per task. Haiku 5.5 uses about 17,000 at low, 33,000 at medium and 162,000 at max. On output alone, in the low tier, medium costs 33,000 × $0.50 per million against 18,000 × $5, about 82% less; max costs 162,000 × $0.50, about 10% less, and if those prompts were over 100k it is 162,000 × $2.50, several times more. Artificial Analysis also put Haiku 5.5 at max at "~3x GPT-6 Luna (max, ~50k)" and noted that at a similar score of 38, Haiku 5.5 at high used about 55,000 tokens to Luna's 50,000. Its cost-per-task figures were provisional on launch day, because its site did not yet model the tiered pricing.
The system card's FrontierCode chart puts output tokens on the x-axis too, and it shows the same knee.

FrontierCode is a nice illustration of a different point. The table compares Haiku 5.5's 46.4% at max with Sonnet 5.5's 52.1% at xhigh, Sonnet's best setting. At max Sonnet 5.5 scored 46.2%, basically tied with Haiku 5.5, after spending something like 580,000 output tokens per task. More effort is not monotonically better on every benchmark, for either model.
Benchmarks, and who ran them
The table compares Haiku 5.5 with Haiku 4.5, GPT-6 Luna and, "for reference," Sonnet 5.5. The system card's summary adds SWE-bench (Pro 64.8%, Multilingual 83.7% against Haiku 4.5's 67.4%, Multimodal 30.7% against 19.8%) and HealthBench Professional (64.8% against 32.2%). Who ran each number matters more than usual, because several are third-party:
- GDPval-AA and AA-Briefcase were run by Artificial Analysis, independently, per the system card. Haiku 5.5 scored 1620 and 1578 Elo against Haiku 4.5's 735 and 614. At medium it scored 1277 on GDPval "while using about a tenth of the output tokens it used at max," and 1372 on AA-Briefcase using "under a quarter."
- OSWorld's GPT-6 Luna figure, 48.9%, was "run by Anthropic on the same 82 tasks through OpenAI's API." Haiku 5.5's 72.4% is the partial-credit score; its strict pass rate, every checkpoint satisfied, is 37.1%.
- Terminal-Bench 4.0: Anthropic reports 39.2% at max in Claude Code. Artificial Analysis's own run got 33%, and the public leaderboard has GPT-6 Luna at 16.4%. 1.8% of Haiku 5.5's trials (12 of 660) were stopped by safety classifiers, which have no fallback model on Haiku 5.5, and all of them failed.
- FrontierCode was run by Cognition, and Cognition's leaderboard has no Haiku 4.5 result, hence the dash.
Artificial Analysis's launch-day summary is the most useful outside view. It puts Haiku 5.5 at 43 on its Intelligence Index at max effort, up from 17 for Haiku 4.5, ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38), and 13 points behind Sonnet 5.5 (56). It also flagged a weak spot: AA-Omniscience knowledge accuracy of 36%, against 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna, with a lower hallucination rate (40% against 55% and 77%). A small model that knows less and says so more often is a reasonable trade for the jobs Haiku is sold for.
Two results surprised me. On ProgramBench, a long-context coding benchmark whose episodes run up to the full 1M window, Haiku 5.5 scored 82.0% against Sonnet 5.5's 79.7%. And Anthropic is unusually blunt about the opposite case: "Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0," where Sonnet 5.5 scored 70.6%.
Speed: "fastest" without a tokens-per-second number
The announcement calls Haiku 5.5 "our fastest model to date," and the footnote narrows it to "at each model's standard speed, although it runs less quickly than our Opus models in Fast Mode." Anthropic publishes no tokens-per-second figure; the models overview just lists its comparative latency as "Fastest."
The measured numbers come from Artificial Analysis, a day after launch: 137 output tokens per second at medium, 160 at high, 178 at low, 187 at xhigh and 242 at max, counted after the first chunk arrives. Its pages show about 90 for Haiku 4.5 and 101 to 129 for Sonnet 5.5. Two adjustments before reading that as "50% faster." The tokens are new-tokenizer tokens, so 137 of them carry roughly as much text as 105 of Haiku 4.5's. And with thinking on, the time you wait is mostly thinking: Artificial Analysis lists 9.93 seconds to the first answer token at low effort. Why max effort streams fastest I don't know; it is their measurement, one day old.
The customer quotes are all relative and unverifiable from outside. Asana reported "over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn" against "the model we use today," which it does not name. Box reported Haiku 5.5 "scored 11 points higher than Haiku 4.5 at about half the latency." If you need latency, measure time to first visible token at the effort you will actually run; the effort setting moves it more than the model choice does.
Context, limits and the migration
The spec sheet, from the Haiku 5.5 model page in the docs:
| Haiku 4.5 | Haiku 5.5 | |
|---|---|---|
| Model ID | claude-haiku-4-5-20251001 | claude-haiku-5-5 |
| Context window | 200k | 1M |
| Max output | 64k | 128k (300k on the Batch API with a beta header) |
| Thinking | manual budget | adaptive, on by default |
| Default effort | none | medium |
| Reliable knowledge cutoff | February 2025 | June 2026 |
| Retirement | not sooner than October 15, 2026 | not sooner than October 7, 2027 |
The retirement row is the one to read twice. Haiku 4.5's commitment runs only to October 15, 2026, a week after this launch, so the migration is not optional for long.
There is a system card, 144 pages, and the capability numbers above come from its
section 8. The migration guide lists the breaking changes, and three of them touch
cost or correctness directly. Thinking counts toward max_tokens, so a small
max_tokens sized for Haiku 4.5 can stop after a thinking block with no text.
temperature, top_p and top_k must be left at defaults or the request returns
a 400. Assistant prefill returns a 400 too, so classifiers that prefilled a
label need structured outputs or a tool with an enum instead. Responses can also
end with stop_reason: "refusal" from safety classifiers, with no server-side
fallback, and Priority Tier is not supported.
The prompting guide has one more cost note worth knowing: in long agent prompts,
"moving from low to medium effort roughly halved early stopping. It also more
than doubled the output tokens for each attempt."
The other price cut
The thread ended with "one more thing": Sonnet 5.5's cache reads drop from $0.20 to $0.10 per million, which Anthropic says makes it "around 20% cheaper to run on most long-running work." The Terminal-Bench chart plots Sonnet 5.5 at both prices, so I decoded that too. Per attempt, the new price is 21% to 26% cheaper across the five effort levels, about $14.11 to $10.45 at max. For that drop, cache reads must have been roughly 40% to 50% of Sonnet 5.5's spend on those runs, which is about what you would expect from an agent rereading its own transcript every turn. The claim holds on Anthropic's own chart.
27.8 billion tokens for $20?
A day after launch, @1kartikkabadi1 posted that "you can get ~27.8b tokens on the $20 plan on Haiku 5.5" and linked a page that works it out. My first reaction was that the arithmetic cannot come from the API. Twenty dollars at Haiku 5.5's cheapest price, the $0.01 cache read, buys 2 billion tokens. Getting to 27.8 billion needs something other than the price list, so I read the page.
It is careful, and it shows its work. It treats a Pro subscription as a fixed monthly pool of usage valued at API list prices, and divides that pool by a blended price per million tokens. The pool is $1,280, which comes from a public project that reads usage panels from saturated Pro accounts: $3.20 of list-price Opus 5.5 usage per percentage point of the weekly meter, times 100 points, times four weeks. The blend assumes 97% of tokens are cache reads, 2.5% are 5-minute cache writes and 0.5% are output, a mix the page takes from measured agentic-coding sessions. On Haiku 5.5's short-prompt prices that is dollars per million, and $1,280 divided by it is 83.5 billion tokens. At the long-prompt prices the blend is five times higher and the answer is 16.7 billion. The 27.8 is neither: it is a 50/50 mix, $1,280 over the average of the two blends. Every price the page uses matches the table above.

So the post's number is $20 multiplied up 64 times by a pool estimate, then divided by a price that is 97% cache reads. Both of those deserve a sentence.
The pool was measured on Opus 5.5. Applying it to Haiku assumes the plan's meter charges each model in proportion to its API list price, so that a percentage point is $3.20 of Haiku just as it is $3.20 of Opus. Anthropic publishes no token allowances, and nobody had a saturated Haiku 5.5 week to measure on the day it shipped. If the meter weighs Haiku more heavily than its price, every Haiku row shrinks, and I have no way to tell from outside. The page itself also leads with the lower of two pool measurements; the other, from Opus 5 accounts in September, is 55% higher. One reply to the post reports 2.2 to 2.5 billion Opus 5.5 tokens a week on a single account, against the page's 0.76 billion. That is one person, but it says the spread is wide.
The cache-read share decides what a "token" is here. Of the 27.8 billion, 27.0 billion are cache reads, the same context re-sent on every turn of an agent loop and billed at a tenth of the input price. Only 139 million are output. That is a fair description of how Claude Code burns tokens, and a misleading one if you read 27.8 billion as text the model writes or reads fresh. A classifier with short, uncached prompts has a blend nearer the $0.10 input price, which turns the same pool into about 13 billion tokens, and nobody runs a classifier through a chat subscription anyway.
Then there is the cliff from earlier. For a workload that is 97% cache reads, the question I could not settle, whether cached tokens count toward the 100,000, is the whole answer: if they do, any agent whose context grows past 100k pays the long-prompt rate on every turn, and the right row is 16.7, not 83.5. The poster's replies say he runs Claude Code with autocompaction at 100k, which is exactly the habit that keeps a session in the cheap tier under the conservative reading. And the 50/50 row splits tokens, not money: at five times the price, the long half eats five-sixths of the pool.
What survives is the ratio. With one pool for every model, Haiku 5.5 at short prompts comes out about 27 times Opus 5.5, because its blended price is about a twenty-seventh of Opus's. That is a real statement about relative cost on the subscription, if the meter follows list prices, and it is the same statement the price table already makes. The absolute 27.8 billion is a ceiling stacked on two estimates, the pool and the traffic mix, plus an unpublished assumption about how Haiku is metered. I would plan around the ratio and treat the headline as an upper bound I have not seen anyone hit.
What I'd do with it
For classification, extraction, routing and short summaries, the migration is close to free money: about 87% off at the same text, as long as you keep prompts under the threshold, pick low or medium effort, and turn thinking off where the task does not need it (it can be disabled at high effort and below). For anything with long prompts, recount in Haiku 5.5 tokens first. Documents that land between roughly 77,000 and 100,000 Haiku 4.5 tokens are where the tier boundary bites, and chunking them is the biggest single saving available. For agent loops, the price is not the variable; effort is. Medium is where the "75% less" lives. Max is where the benchmark table lives, and on most of Anthropic's own charts it costs more than the model it replaces.
The related pieces on this site go at the same question from other sides: the harness effect on why orchestration sets an agent's token bill more than the model does, how LLM inference works for why prefill and decode cost what they do, and tiny browser models, which priced a Haiku 4.5 call against a 28 KB model; at Haiku 5.5's prices that gap shrinks about tenfold.
How I checked
I read the announcement page, its footnotes and the X thread from @claudeai through
the fxtwitter mirror, the 144-page system card (sections 1, 2 and 8 in full), and
the docs pages for the models overview, Haiku 5.5 overview, what's new, migration
guide, pricing, context windows, effort, prompt caching and the Haiku 5.5 prompting
guide, all on October 8. The per-task costs come from decoding the four inline-SVG
charts on the announcement page: I took each axis's tick positions, fitted a log
scale for cost and a linear one for score, and mapped every <circle> back to data.
I checked the decode against the system card's own numbers (GDPval medium 1275
against 1277, OSWorld and Terminal-Bench max exact). Workload costs are list prices
times the token counts stated, with a 1.3x tokenizer factor from the docs; that factor
is Anthropic's approximation, and the real one depends on your text. Speed figures
are Artificial Analysis's, read from its model pages, and I did not measure latency
myself. I did not call the API for this piece, so the threshold question, whether
cached tokens count toward 100,000, is unverified. For the subscription estimate I read
the post and its replies through the fxtwitter mirror and rendered the lineup page in a
headless browser; its blends and token counts recompute exactly from its stated formula
and prices, and its $1,280 pool is the page's own sourced estimate, which I could not check.