NotCheapMAKE EVERY CREDIT COUNT

Firecrawl vs Jina AI Reader vs Diffbot: Real Web Extraction Cost per 1 Million Pages in 2026

The short answer: the three price extraction in fundamentally different units, and the harder the page mix, the more that difference matters. Firecrawl and Diffbot charge a fixed number of credits per page (with more credits for tougher extraction); Jina AI Reader charges by the tokens in the output it returns, so a long page costs proportionally more regardless of how hard it was to fetch. On easy static pages, all three land near a fraction of a cent per page. On difficult ecommerce, news or social pages with anti-bot defenses, where the ILLUSTRATIVE success rate drops to 60%, the cost per attempted page rises to roughly $0.0042 (Firecrawl), $0.0024 (Jina), and $0.01 (Diffbot). Firecrawl's and Diffbot's cost per successfully extracted page is 40–67% higher than their per-attempted-page cost, because a failed attempt still consumes a credit and must be retried; Jina's does not carry this penalty at all, because Jina does not charge for failed requests in the first place, so its per-successful-page cost equals its per-attempted-page cost divided by the success rate that already went into computing the attempted figure — in effect, Jina's billed cost tracks successful extractions directly.

What each vendor bills

Firecrawl (OFFICIAL, firecrawl.dev/pricing). Credit-based: Free (1,000 credits/month), Hobby ($16/mo annual, 5,000 credits), Standard ($83/mo, 100,000 credits), Growth ($333/mo, 500,000 credits), Scale ($599/mo, 1,000,000 credits). One credit is one basic scrape, crawl page, or map page; JavaScript rendering is automatic and included at no extra credit cost. Search costs 2 credits per 10 results. AI-prompted extraction with a JSON schema was reported at roughly 5–10 credits per URL, and JSON-mode change-tracking comparison costs 5 extra credits per page. At the Standard plan, the effective rate is about $0.00083 per credit; at Growth, about $0.00067. Credits do not roll over on self-serve plans.

Jina AI Reader (OFFICIAL for the mechanism, REPORTED for the rate). Token-metered, not page-metered: Reader (r.jina.ai) converts a URL to markdown and bills for the tokens in the output it returns, at a reported rate of roughly $0.045–$0.05 per million tokens on paid top-ups (one separate source cites $0.02 per 1,000 tokens, equivalent to $20 per million, an order of magnitude higher than the other figure; this article uses the lower, more recently sourced $0.045–$0.05 range and flags the conflict). Every new API key includes 10 million free tokens. Failed requests are not charged, which means Jina's billing model is structurally different from Firecrawl's and Diffbot's: a page that fails to extract costs Jina nothing at all, so Jina's cost tracks successful extractions directly rather than every attempt. Using ReaderLM-v2 (the model-based extraction mode for harder pages) roughly triples the token cost. Jina AI was acquired by Elastic in October 2025, adding enterprise distribution to the product (REPORTED); this does not change the pricing mechanism as documented.

Diffbot (OFFICIAL, diffbot.com/pricing). Credit-based: Free ($0, 10,000 credits/month, 5 calls/minute), Startup ($299/mo, 250,000 credits at $0.001 each, 5 calls/second), Plus ($899/mo, 1,000,000 credits at $0.0009 each, 25 calls/second, includes Crawl), Enterprise (custom). Extract 1 page = 1 credit; extract with a datacenter proxy = 2 credits. Knowledge Graph activities are billed far more heavily (25 credits per entity export, 100 credits per facet query or entity refresh) and are a different product from page extraction. Crawl is gated to Plus and above.

Three page mixes and a success-rate model

The page mixes and success rates are ILLUSTRATIVE. A: static pages (plain HTML, no anti-bot defenses, 95% success rate, ~2,000 tokens per page for the Jina calculation). B: JS-heavy SaaS or product pages (requires rendering and sometimes a proxy, 85% success rate, ~8,000 tokens per page). C: difficult ecommerce, news or social pages with retries (anti-bot defenses, CAPTCHAs, or aggressive rate limiting, 60% success rate, treated as needing Jina's ReaderLM-v2 mode at ~90,000 effective tokens per page after the 3x multiplier).

Per-page credit or token assumptions: Firecrawl 1 credit (A, B), 5 credits (C, reflecting AI-prompted extraction with retries), charged on every attempt regardless of success. Diffbot 1 credit (A), 2 credits (B, datacenter proxy), 10 credits (C, ILLUSTRATIVE for repeated attempts against anti-bot defenses; Diffbot's own published rate table does not itemize a specific "difficult page" tier, so this is a modeling assumption, not a vendor figure), also charged on every attempt. Jina token counts as above, charged only on the attempts that succeed.

Cost per 1 million attempted pages

Formula: Firecrawl/Diffbot cost = attempted pages × credits per page × price per credit (using Firecrawl's Standard rate of $0.00083/credit and Diffbot's Startup rate of $0.001/credit), charged whether or not the attempt succeeds; Jina cost = attempted pages × success rate × tokens per page × $0.045 ÷ 1,000,000, because only the successful attempts are billed.

Page mixFirecrawlJina AI ReaderDiffbot
A: static$830$85.50$1,000
B: JS-heavy$830$306.00$2,000
C: difficult, with retries$4,150$2,430.00$10,000

At 10,000 attempted pages, these scale down to $0.86–$10 (A), $3.06–$20 (B), and $24.30–$100 (C); at 10 million attempted pages, $24,300–$100,000 (Jina to Diffbot) on the hardest mix, with Firecrawl at $41,500.

Cost per 1 million successfully extracted pages

This is the number that actually matters, because a failed attempt on a difficult page is not a usable result. Formula: cost per successful page = cost per attempted page ÷ success rate.

Page mixSuccess rateFirecrawl, per successful pageJina, per successful pageDiffbot, per successful page
A: static95%$0.00087$0.00009$0.00105
B: JS-heavy85%$0.00098$0.00036$0.00235
C: difficult, with retries60%$0.00692$0.00405$0.01667

Retry/failure amplification, the ratio of cost-per-successful to cost-per-attempted, applies only to Firecrawl and Diffbot, since both charge a credit on every attempt: 1.05x at 95% success, 1.18x at 85%, and 1.67x at 60%. Jina carries no such amplification: because it never bills a failed request, its cost-per-successful-page is identical to what its cost-per-attempted-page would be if every attempt succeeded, at every success rate. On the hardest page mix, Firecrawl and Diffbot spend nearly two-fifths of their attempted-page budget on attempts that ultimately fail and must be retried; Jina spends nothing on those failures at all.

Attempted pages are not successful pages

Never conflate the two. At 1 million attempted difficult-mix pages, Firecrawl costs $4,150 and produces roughly 600,000 successful extractions ($0.00692 each), while Jina, billing only the successful subset, costs $2,430 for that same 600,000 successful extractions. To reach 1 million successful extractions on that same mix requires about 1.67 million attempts on Firecrawl and Diffbot, costing roughly $6,917 on Firecrawl and $16,667 on Diffbot; Jina's cost for 1 million successful extractions is simply $4,050 (1,000,000 × $0.00405), since its billing already tracks successes rather than attempts. Any vendor quote or internal budget that states "cost per million pages" without specifying attempted versus successful is describing two different numbers that can differ by 5–67% depending on page difficulty.

Sensitivity

  1. Which Jina rate is real. The $0.045–$0.05 per million tokens figure versus the $20 per million figure from a separate source is a 400x difference; confirm the current rate directly before budgeting Jina at scale.
  2. Page size on Jina. Because Jina meters output tokens, a documentation page rendering to 40,000 tokens costs 40 times a landing page rendering to 1,000 tokens, an input the caller does not fully control in advance.
  3. Success rate. On Firecrawl and Diffbot, moving from 85% to 60% success on the same page type raises cost per successful page by roughly 70% (the 1.18x to 1.67x amplification difference), because both bill every attempt. Jina's cost per successful page does not move at all with success rate, since it only bills successes in the first place; only the page's token count changes Jina's per-successful-page cost.
  4. Diffbot's Knowledge Graph meter. Entity exports and facet queries cost 25–100 credits, 25–100 times a plain page extraction; conflating page-extraction volume with Knowledge Graph usage badly understates a Diffbot bill that uses both.

Budgeting traps

  • Quoting "cost per page" without saying attempted or successful. These differ by the retry amplification factor, which grows sharply on harder page mixes.
  • Assuming JS rendering costs extra everywhere. Firecrawl includes it in the base credit; other vendors in the broader market charge 2x or more for it, so this assumption should be confirmed per vendor rather than applied uniformly.
  • Token-based billing without a size estimate. Jina's bill is unknowable in advance without an estimate of average output token count across your specific page mix.
  • Treating Knowledge Graph credits like extraction credits. On Diffbot they are a completely different order of magnitude.

What to ask before you buy

Ask Jina for its current confirmed per-million-token rate in writing, since third-party sources disagree by two orders of magnitude. Ask Firecrawl and Diffbot for the exact credit cost of your hardest page type, including any proxy or AI-extraction surcharge, and measure your actual success rate on a sample of 100–500 real target pages before projecting cost at scale.


ARTICLE 10