~/satyajit

AnyAPI's Scrape router: a ladder of scrapers, priced by the rung that gets through

mdjsonmcp

2026-10-06 · 17 min · agents · tool-calling · benchmarks · pricing · browser-automation

Why read this

Notabletop 60%

Recounts AnyAPI's 100-page scrape benchmark from its CSV, reprices Firecrawl by plan, and replays a router built from the rivals: 95/100 at about $1.23.

  • Analysis found nowhere else
  • Concrete numbers to act on
  • Explained from first principles

Developer tools & infraAPI onlyProprietaryPractitioner tool

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
1 of 3: API-only, gated or restrictive licence
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
1 of 3: A specialist community
Only here?
3 of 3: The only place this analysis exists

Score 59 of 100, ranked 248 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Kevin Wang posted the launch of AnyAPI's Scrape router on X with an 11.7-second film and two lines of numbers: on 100 easy, medium and hard pages, AnyAPI got 98% at $1.05 per thousand, and Firecrawl on its Hobby plan got 90% at $3.97. The second post in the thread linked a methodology gist with a per-page CSV.

I looked because I had just read treg, which does the same thing for people search: put a dozen paid providers behind one endpoint, try the cheap ones first, and only pay for the one that answers. Scraping is a better fit for that idea than email lookup. Whether a page comes back depends far more on which site you hit than on which vendor you pay, and the vendors all fail on different sites.

What I found, briefly. The benchmark is the founder's own (the gist is kev1n, and AnyAPI's author page lists that GitHub account for Kevin Wang, "Founder, AnyAPI"). Every number in its table reproduces from its CSV. The success gap is real but narrow, and the price gap mostly comes from which Firecrawl plan you price. The more interesting result is one the gist doesn't draw: if you replay its own per-page outcomes through a router built from the four rival services, you get most of AnyAPI's result. The routing is what earns the 98.

A chart titled '100 web pages, 40 sites, one try each'. Five rows of 100 thin bars, shaded easy, medium and hard. AnyAPI 98/100 at $1.05 per 1,000 pages; Firecrawl 90/100 at $3.97; Bright Data 81/100 at $1.50; Jina Reader 55/100 at $0.73; Cloudflare Browser Rendering 48/100 at $0.10 plus a $5 a month plan.
The launch chart. Each row's bars are sorted by tier, not lined up by page, so the chart cannot show which pages each service missed; the CSV can. The price column is per 1,000 real pages, not per 1,000 requests (frame from the launch film on X, @mxfp4).

One request, five lanes

The product is a single endpoint, POST /v1/run/web.scrape. You send a URL and get back Markdown, HTML, a title and a description. Behind it sit five providers that AnyAPI names only by animal: Bison, Heron, Salamander, Koala and Dolphin. The public catalog at api.getanyapi.com/catalog lists them in call order with a flat price each, and the scrape page renders the same table.

AnyAPI's web scrape pricing table, headed 'one API, 5 interchangeable providers. The cheapest provider serves first. If it fails, the next one takes over in the same call.' Rows: 1 Bison, serves first, 71.8% of calls, 58,240 calls, 97.67% uptime, 4.10% blocked, 2495 ms, $0.70 per 1k; 2 Heron, 0.8%, 92.55%, $1.00; 3 Salamander, 21.4%, 90.95%, $1.25; 4 Koala, 3.9%, 62.97%, $1.50; 5 Dolphin, 2.1%, 77.80%, $1.50.
The five lanes, cheapest first, with 30-day share, uptime, blocked rate and median latency. Rendered by me in headless Chromium on 2026-10-06; the note under the table says these are not head-to-head numbers, because each lane only sees what the lanes above it failed (getanyapi.com/api/web/scrape).

Bison costs $0.70 per thousand and takes 71.8% of the traffic. Heron is $1.00, Salamander $1.25, Koala and Dolphin $1.50. A call starts on Bison. If Bison fails, the same call goes to Heron, and so on down. You are charged at the price of the lane that served you. The launch film shows this on a Zillow listing: Bison gets a 403, the card reads "Blocked, not charged", and the next lane returns the house.

A frame from the launch film. On the left, a list of lanes with prices: Bison $0.70, Salamander $1.00, Heron $1.00, Koala $1.50. A line runs from Bison to a card on the right reading '403, Blocked, not charged'.
The first lane is refused by the site and the attempt is not billed. Note the lane order and prices in the film (Salamander second at $1.00) differ from today's catalog (Heron second; Salamander $1.25) (launch film, @mxfp4).
A later frame. The $1.00 lane is highlighted, and the card now shows the scraped Markdown: $9,500,000, # 2700 Point Ln, Highland Park, IL 60035, 9 beds, 19 baths, 32,683 sqft, Built in 1995, ## Facts & features.
The second lane gets the page, and that lane's price is the charge for the call (launch film, @mxfp4).

A few details in the API reference matter more than the film suggests.

The output schema carries x-anyapi-empty-fails-over: true. An empty answer counts as a failure and moves the call down the ladder, which is the right default for scraping: the classic soft block is an HTTP 200 with a captcha or an empty JavaScript shell. Whether the gateway catches a 200 that contains a real-looking block page, I couldn't tell from the docs. The benchmark's own scorer does catch those, by checking for the expected title or text, so the gist's numbers are not inflated by it either way.

The catalog entry carries a ceiling, failoverMaxPer1kUsd: 1.5. However far down the ladder a call goes, it costs at most $1.50 per thousand. A max_cost_usd query parameter lets you lower that per call; lanes above your cap are skipped, and if none is left the call is refused with nothing charged. You can also pin lanes with source and allowFallbacks: false, skip them with ignoreSources, or ask for speed with preferLatencyUnderMs, which picks the cheapest lane whose 30-day median is under your number and charges its price.

One sentence in the FAQ softens "not charged": "Some failed wallet-funded requests incur processing charges that we pass through at cost". The film's 403 is not billed; I can't say which failures are.

AnyAPI's MIT-licensed CLI is the only code it publishes. It is a client. The router itself runs server-side and is closed. The CLI does confirm one small thing about pricing: "one internal credit is $0.00001 and lane prices are whole credits" (src/format.ts:16), so every lane price is a whole number of cents per thousand calls.

The ladder inside the ladder

When people say "scrape router" they usually mean an escalation ladder over techniques, and it helps to have that picture first.

The bottom rung is a plain HTTP fetch from a datacenter IP: a GET with believable headers. It costs almost nothing and gets Wikipedia, docs sites and most news. The next rung is a headless browser, for pages whose content only exists after JavaScript runs. That costs CPU-seconds and memory per page. The rung after that changes where the request comes from: residential or mobile proxy IPs, billed by bandwidth, because the big anti-bot services score the IP's reputation before they look at anything else. The top rungs are the "unblocker" products that bundle all of it with browser fingerprint management and captcha solving, priced per successful request. Each rung is slower and dearer than the one below, and each one gets through on sites where the one below gets refused.

Firecrawl's own docs show a two-rung version of this inside one vendor. Its auto proxy mode starts on basic proxies and, "If the target responds with 401, 403, or 429", retries the same URL on enhanced proxies. Other codes, including 404 and 5xx, do not escalate, "because a different proxy would not change the answer". That rule is worth stealing. A router should escalate on evidence that the requester was the problem, not on any failure.

AnyAPI does not run that ladder. Its FAQ says so directly: "Each source behind an endpoint handles its own rendering, proxying, and unblocking", and "AnyAPI does not sell proxy bandwidth". The router is a ladder of vendors, each of which is itself a ladder of techniques that you cannot see. Bison, at $0.70, is presumably a cheap rung-one-or-two service that gets most pages. Koala and Dolphin at $1.50 are priced like unblockers. That is my inference from the prices. AnyAPI does not say who or what the animals are.

I think this is the right layer for a small company to work at. Running your own residential proxy pool is a capital-and-abuse-desk business. Buying five services that already did, and ordering them, is a spreadsheet problem with a retry loop. The cost is that you inherit each vendor's latency on every rung you fall through: Bison's 30-day median is 2,495 ms and Heron's is 4,254 ms, so a call that falls to the second lane waits roughly for both.

What a page costs when failures are free

Treg's router sorts candidates by an expected cost per hit,

cost per hit=price×P(billed)P(hit),\text{cost per hit} = \frac{\text{price} \times P(\text{billed})}{P(\text{hit})},

and the treg article pointed out that when a provider bills only on success, P(billed)=P(hit)P(\text{billed}) = P(\text{hit}) and the expression collapses to the price. AnyAPI's lanes are in that regime: a blocked attempt is free, and you pay the flat price of the one lane that served you. So the cost of a page is simply the price of the first lane that gets it, and the cost-optimal order is the one AnyAPI uses: cheapest first. The hit rates don't enter the ordering. They only decide how often you fall through.

That changes as soon as a failed attempt costs something. For a ladder where lane ii costs cic_i per attempt whether or not it works, and gets the page with probability pip_i, the expected cost of trying it is cic_i and the expected yield is pip_i, and the order that minimises expected spend per page is ascending ci/pic_i / p_i. That is treg's formula again with P(billed)=1P(\text{billed}) = 1. A cheap lane that rarely works should sit below a dearer one that usually does. You will see this in the replay further down, where Firecrawl's per-call credits push it behind Bright Data's pay-on-success price at one plan rate and ahead of it at another.

The benchmark's $1.05 tells you something about how far its pages fell. If every one of the 98 successes had been served by Bison, the run would have cost $0.70 per thousand. A call cannot cost more than $1.50. Taking the average at face value and the failed attempts as free, at least 43 of the 98 pages must have been served below Bison:

98×0.70+e×(1.50−0.70)≥98×1.045  ⇒  e≥43.98 \times 0.70 + e \times (1.50 - 0.70) \ge 98 \times 1.045 \;\Rightarrow\; e \ge 43.

So on this page list, at least 44% of successes needed a backup lane. That is the benchmark being harder than real traffic, which is fine; it was built to be. AnyAPI's own usage page puts production web.scrape at 86,007 requests for $68.28 over 30 days, about $0.79 per thousand, with 9.7% of calls rescued by a later lane. Its home page reports the same pattern for its five busiest endpoints over 86 days: the first provider returned 92.2% of 96,213 calls on its own, AnyAPI returned 99.6%, and rescuing 7,115 of the 7,472 misses "added 4.1% to the bill across those calls". Failover is cheap when most pages are easy, and the price you see depends heavily on your mix.

Recounting the benchmark

The gist holds two files: a README with the method and a 100-row table, and results.csv with one row per page and an ok or failed per service. No harness code. I recounted the CSV and checked it row by row against the README table; they agree on every cell, and every total in the headline table reproduces: AnyAPI 98 (30 easy, 32 medium, 36 hard), Firecrawl 90 (30, 31, 29), Bright Data 81 (28, 28, 25), Jina 55 (30, 20, 5), Cloudflare 48 (30, 12, 6).

The page list is reasonable. Thirty easy pages from Wikipedia, MDN, the Python and Rust docs, arXiv, Gutenberg, Hacker News and three news sites. Thirty medium pages from retail, travel and listings (Amazon, Walmart, Zillow, Target, Best Buy, IMDb, Booking, Airbnb, eBay, Costco), three per site. Thirty hard pages, three per site, from Shein, G2, Hyatt, Lowe's, leboncoin, Tripadvisor, Etsy, idealista, realtor.com and Yelp. Then ten single pages taken from Proxyway's 2025 scraping API report, eight of them filed as hard and Google search and YouTube as medium. That is 40 sites. The success test is the right one: the expected title or text must be present and no bot wall found, and a block page served with HTTP 200 counts as a failure. One request per page per service, measured 2026-09-29 to 2026-10-01.

Three things in the CSV are worth more than the headline.

AnyAPI's successes are a superset of everyone else's. There is no page that any rival got and AnyAPI missed. Its only two failures are two Shein pages that every service failed. If you take the union of the four rivals, they get 95 pages between them, and AnyAPI's extra three are all Yelp.

The lead over Firecrawl is eight pages, and they cluster. Firecrawl failed 10 pages on 6 sites: all three Shein pages, all three Yelp pages, and Google search, ImmobilienScout24, Instagram and Nordstrom. AnyAPI got eight of those. Counted by site instead of by page, AnyAPI beat Firecrawl on 6 of the 40 sites and lost on none. That is a consistent edge, and a one-sided sign test over the 8 discordant pages gives about 0.004, but the 3 Yelp pages are one decision by one anti-bot vendor on one day. Swap Yelp for another hard site and the gap could be five pages or ten.

The measurement is one try per page. That is fair to both sides in one sense: AnyAPI's single call contains up to five attempts, and Firecrawl's enhanced mode contains its own retries. It is less fair in another. A page that fails once might pass a second time, and the gist does not say how often that happens for any service. Latency is not reported at all, and the router is the service most likely to be slow on a hard page, since it pays for every lane it falls through.

The Hobby plan

The headline price comparison is $1.05 against $3.97. The gist says where $3.97 comes from: "Firecrawl credits at the Hobby plan rate ($19 for 5,000 credits; larger plans cost less per credit)". That parenthetical carries most of the gap.

Working backwards, $3.97 per thousand real pages on 90 real pages is a run cost of $0.357, which at $0.0038 a credit is about 94 credits for 100 calls. Firecrawl's docs say enhanced proxies cost "1 credit per request", the same as basic, so the run billed a little under one credit per call. Now reprice those 94 credits on Firecrawl's pricing page as I read it on 2026-10-06. Hobby billed yearly is $16 for 5,000 credits, which gives $3.34 per thousand real pages. Standard billed yearly is $83 for 100,000 credits, which gives about $0.87. Growth, at $333 for 500,000, gives about $0.70. On a per-page basis at any plan above Hobby, Firecrawl on this benchmark is as cheap as AnyAPI or cheaper. AnyAPI's own home page uses Firecrawl's yearly Hobby rate, $3.20 per thousand, for its general price comparison, so neither the gist nor the home page hides the basis. The X post just leaves it out.

Per-page price is still the wrong comparison, because a plan is a monthly commitment and a wallet is not. At the benchmark's mix:

A month of scraping: wallet or plan?
AnyAPI wallet · 98/100 on the bench$10.50/mo
Firecrawl, Hobby · 90/100 on the bench$46.00/mo
Bright Data · 81/100 on the bench$15.00/mo
All three priced at the benchmark's page mix: AnyAPI at $1.05 and Bright Data at $1.50 per 1,000 real pages, Firecrawl at about 94 credits per 90 real pages on the cheapest plan plus $5 credit packs that covers the month (yearly-billed prices from firecrawl.dev/pricing, 2026-10-06). The success rates differ, so the services do not deliver the same pages; a page Firecrawl missed on the bench is not bought back by spending more.

Below about 1,000 real pages a month, Firecrawl's free tier wins outright. From there to about 79,000 pages a month, AnyAPI's wallet is cheaper than any Firecrawl plan, because you are paying $16 or $83 for credits you don't use. Between about 79,000 and 105,000, Standard's 100,000 included credits make Firecrawl cheaper. Past them, Standard's top-ups cost $5 per 2,000 credits, dearer per page than AnyAPI, so the wallet wins again until about 317,000 pages, where Growth's $333 plan takes over and Firecrawl stays cheaper from there up. Bright Data, at $1.50 per success, sits above AnyAPI across the whole range. None of this changes the success rates: on this list you get about 8 more real pages per hundred from AnyAPI whatever you pay Firecrawl.

So the honest version of the headline is something like "more pages than Firecrawl on this list, and cheaper unless you scrape enough to fill a Standard or Growth plan". That is still a good pitch for anyone whose scraping is bursty or small, which describes most agent workloads.

Building the router out of the rivals

This is the test I wanted. The CSV records, for each page, which of the four rival services got it. That is enough to replay a ladder built from them: put them in an order, send each page down until one of them got it, and charge each attempt what that service charges.

The gist publishes only a cost per real page, so I back out an average cost per call. Cloudflare comes to about $0.05 per thousand calls (plus its $5 a month plan, which I leave out), Jina to about $0.40, Firecrawl at the Hobby rate to about $3.57. Bright Data charges only on success, $1.50 per thousand, so its failures are free. These are averages; a long page costs Jina more tokens and a slow one costs Cloudflare more browser time.

Build the router out of the rivals

Order the rungs, switch them on or off. Each page goes down the ladder until a rung returned the real page in the benchmark.

  1. 148 served
  2. 212 served
  3. 325 served
  4. 410 served
easy 1-30medium 31-60hard 61-100 (93, 97 medium)
Your ladder
95/100 at $1.23
per 1,000 real pages
AnyAPI, as published
98/100 at $1.05
per 1,000 real pages
Outcomes are the gist's own per-page results. Per-attempt costs are averages back-derived from each service's published cost per real page, so they ignore that a long page costs Jina more tokens and a slow one costs Cloudflare more browser time; Cloudflare's $5 monthly plan and every subscription commitment are left out. Dashed cells: no rung in your ladder got the page.

Cheapest first, Cloudflare, Jina, Bright Data, Firecrawl: 95 of 100 at about $1.23 per thousand real pages. Cloudflare serves the 48 pages it can, Jina picks up 12 more, Bright Data 25, and Firecrawl 10. Dropping Jina barely moves it: Cloudflare, Bright Data, Firecrawl gets 95 at about $1.20.

Then tick the Standard-plan box. Firecrawl's per-call cost falls to about $0.78 per thousand, now cheaper than Bright Data's $1.50 per success, and the ci/pic_i / p_i ordering says Firecrawl should move up. Cloudflare, Jina, Firecrawl, Bright Data gets the same 95 pages for about $0.66 per thousand. The order that was right at one price is wrong at another, which is the whole case for a router that reads prices from data rather than from a hard-coded list.

What's left for AnyAPI is the three Yelp pages, the two that nobody got, and the plumbing: one key, one schema, no plans, and a ladder someone else keeps ordered as prices and block rates move. Those are worth something. They are not a scraping technology. On this benchmark, a router over four off-the-shelf services, built from data anyone can download, lands within three pages and about 20 cents of the product.

What I'd use it for

If I needed a few thousand pages a week for an agent, mostly easy with a long tail of retail and listings sites, I would use this or something shaped like it, and I would set max_cost_usd so a bad week can't surprise me. The pay only for what got through model is the part I like most, and it is what makes cheapest-first ordering correct rather than merely cheap.

If I were scraping hundreds of thousands of pages a month on a known set of sites, I would measure my own sites through two or three vendors, buy a plan from the one that wins, and keep a second as a fallback. That is a two-rung version of the same router, and at that volume the plan's included credits are cheaper than any per-request wallet in the table above.

And I would not read much into 98 against 90. It is eight pages on six sites, one try each, over three days, measured by the vendor. The page list and the scoring are good, and the CSV is honest enough that I could find every caveat in it. That is more than most launch benchmarks give you.

How I checked

I read the X post and its thread through the fxtwitter mirror and pulled the launch film's frames with ffmpeg; the four figures here are three film frames and one screenshot of AnyAPI's scrape page, rendered in headless Chromium. I cloned the gist (revision a99fa19, 2026-10-02), recounted results.csv in Python, checked it cell by cell against the README table, and computed the union, the per-site discordance and the sign test from it. The ladder replay uses those per-page outcomes directly; its per-call costs are back-derived from the gist's own prices as described above. AnyAPI's lane prices, health figures and the failoverMaxPer1kUsd ceiling come from the public catalog at api.getanyapi.com/catalog; routing parameters come from the web.scrape API reference; production volume and spend come from AnyAPI's usage page, all read on 2026-10-06. Firecrawl's plan prices and credit rules come from its pricing page and its proxy docs, read the same day. I cloned AnyAPI's CLI (commit fb283ae) to confirm that the router is not in it. I did not call AnyAPI's API or any scraper, so the success rates are the gist's, not mine, and I could not check which providers sit behind the five animal names or which failed calls carry a processing charge.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "AnyAPI's Scrape router: a ladder of scrapers, priced by the rung that gets through", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026anyapiscraperouter,
  author = {Satyajit Ghana},
  title  = {AnyAPI's Scrape router: a ladder of scrapers, priced by the rung that gets through},
  url    = {https://ai.thesatyajit.com/articles/anyapi-scrape-router},
  year   = {2026}
}
share