Mistral OCR vs AWS Textract vs Google Document AI: Real Document AI Cost per 1 Million Pages in 2026
The short answer: Mistral OCR is dramatically cheaper for complex extraction specifically, because it collapses everything beyond basic OCR into one flat "Document AI mode" price instead of the steep per-feature tiering AWS and Google both use. At 1 million pages, basic OCR costs about $1,500 on both AWS Textract and Google Document AI, against $4,000 on Mistral's standard API (or $2,000 on its batch API) — Mistral is actually the most expensive of the three for plain OCR. But for structured extraction (tables, forms, custom fields), the ranking flips hard: Mistral's flat $5,000 Document AI mode covers everything from table extraction to custom fields, while the same complexity costs $15,000 to $65,000 on AWS Textract (depending on whether Tables, Forms, or both are needed) and $10,000 to $30,000 on Google Document AI (Layout Parser versus Custom Extractor). Mistral's pricing has itself moved a great deal: its OCR line went from about $1 per 1,000 pages at launch in March 2025 to $4 per 1,000 for the current OCR 4 generation, a 4x increase as the model added capability across four major versions in fifteen months.
What each vendor bills
Mistral OCR (OFFICIAL, Mistral's own pricing, current OCR 4 generation, independently re-verified by multiple sources through August 2026). Standard API: $4 per 1,000 pages. Batch API: $2 per 1,000 pages (50% off). Document AI mode (full structured extraction — tables, layout, entities): $5 per 1,000 pages, a single flat rate regardless of document complexity, unlike AWS's and Google's steep per-feature tiers. Mistral OCR has iterated fast: the original OCR (March 2025, 11 languages, roughly $1 per 1,000 pages) was deprecated in February 2026 in favor of OCR 3, which was itself succeeded by OCR 4 (mid-2026, 170 languages claimed by some trade coverage though Mistral's own documentation more conservatively states 40-plus). Older third-party guides citing $1–$2 per 1,000 pages for Mistral OCR are describing a prior, now-superseded generation, not the current rate.
AWS Textract (OFFICIAL, AWS Textract pricing). DetectText (basic OCR): $1.50 per 1,000 pages. AnalyzeDocument with Tables: $15 per 1,000. AnalyzeDocument with Forms: $50 per 1,000. Tables and Forms combined: $65 per 1,000. AnalyzeExpense (invoices/receipts): approximately $7 per 1,000. A free tier offers 1,000 pages a month per API for the first 3 months (DetectText and AnalyzeDocument allowances apply independently, so both can be used free in the same month). Textract's pricing structure is the steepest of the three by feature: moving from basic OCR to full Tables-and-Forms extraction is a 43x jump in per-page cost.
Google Document AI (REPORTED, with an unresolved 2.3x conflict on the base OCR rate). OCR/Read API: reported at $1.50 per 1,000 pages in the majority of sources reviewed, though one source states $0.65 per 1,000 — this article could not reconcile the two and flags both. Layout Parser (Gemini-powered layout parsing, released December 2025): $10 per 1,000 pages. Form Parser and Custom Extractor: $30 per 1,000 pages. Independent cost analysis reports that Google Document AI's real total cost for a moderate deployment (around 5,000 pages a month) typically runs 3 to 5 times the listed per-page API price, because Google's specialized processors carry a separate hourly charge (reported around $0.05/hour) on top of the per-page rate — a cost that amortizes down as volume rises but can dominate the bill at lower, steadier usage.
Three page-mix shapes
All ILLUSTRATIVE, matching this site's established document-parsing comparison framework. A: digital/plain-text pages (basic OCR only). B: scanned/table-heavy pages (structured table or layout extraction needed). C: complex forms and custom fields (the most feature-intensive tier each vendor offers).
Cost per 1 million pages
Formula: cost = pages ÷ 1,000 × rate per 1,000.
| Page mix | Mistral OCR | AWS Textract | Google Document AI |
|---|---|---|---|
| A: basic OCR (standard rate) | $4,000 | $1,500 | $1,500 (or $650 at the conflicting lower rate) |
| A: basic OCR (batch, Mistral only) | $2,000 | — | — |
| B: tables/layout | $5,000 (Document AI mode, flat) | $15,000 (Tables only) | $10,000 (Layout Parser) |
| C: forms/custom extraction | $5,000 (same flat Document AI mode) | $50,000–$65,000 (Forms, or Tables+Forms) | $30,000 (Custom Extractor) |
At the complex end (shape C), Mistral's flat rate is 10 to 13 times cheaper than AWS Textract's combined Tables-and-Forms tier, and 6 times cheaper than Google's Custom Extractor — the single largest structural advantage in this comparison, and it comes specifically from Mistral not tiering its price by extraction complexity the way both legacy cloud providers do.
Why basic OCR alone favors AWS and Google
This is the one shape where the legacy providers win outright: at 1 million plain-text pages, AWS Textract and Google Document AI's OCR-only tier both undercut even Mistral's batch rate ($1,500 versus $2,000). A workload that genuinely never needs structured extraction — pure text digitization, search indexing, archival OCR — is cheaper on the legacy providers' basic tier than on Mistral, even before considering the reported hidden hourly-processor multiplier that applies mainly to Google's specialized (not basic OCR) processors.
The self-hosted alternative, for context
Independent benchmarking reports that current open-weight OCR models (GLM-OCR, PaddleOCR-VL) now score above several proprietary vision-language models on the OmniDocBench leaderboard, and that running one of these self-hosted on a rented GPU costs roughly $0.09 per 1,000 pages — a 333x to 722x cost ratio against the managed cloud structured-extraction range this article models across Google (Custom Extractor, ~$30/1,000) and AWS (Tables and Forms, up to ~$65/1,000), not a single vendor's rate, at the cost of operating the GPU infrastructure and inference pipeline yourself rather than calling a managed API. This is not one of this article's three primary comparisons, but it is the honest ceiling on what a team with the engineering capacity to self-host could pay instead.
Sensitivity
- Extraction complexity. The single largest lever for AWS and Google, where moving from basic OCR to full structured extraction multiplies cost by 20–43x; Mistral's flat two-tier structure makes this a non-issue past the initial OCR-versus-Document-AI-mode choice.
- Which Google OCR rate is real. A 2.3x conflict between $0.65 and $1.50 per 1,000 pages across sources reviewed; confirm directly before budgeting basic-OCR volume on Google.
- Deployment volume, for Google specifically. The reported 3–5x real-cost multiplier applies most strongly at moderate, steady volumes (around 5,000 pages/month); it likely amortizes down at the 1-million-page scale modeled here, though an exact scaled figure was not found.
- Batch eligibility on Mistral. A confirmed 50% discount for asynchronous processing, applying only to the standard OCR tier in the sources reviewed, not confirmed for the Document AI mode specifically.
Budgeting traps
- Assuming Mistral is always the cheapest option. It is the most expensive of the three for plain OCR at this volume; its advantage is specifically in structured/complex extraction.
- Pricing a mixed document set at a single tier's rate. A real document set almost never needs the same extraction complexity throughout; blending shapes A, B and C proportionally changes AWS's and Google's totals far more than Mistral's flatter structure.
- Ignoring Google's separate hourly processor charge for specialized extraction. It can push real cost to 3–5x the listed per-page rate at moderate, steady volume.
- Budgeting off an older Mistral OCR rate. The current OCR 4 generation prices roughly 4x higher than the original OCR line from March 2025, which some third-party guides still cite.
What to ask before you buy
Confirm Google's current OCR/Read API rate directly, given the unreconciled 2.3x conflict between sources, and ask specifically about the separate hourly processor charge if using a specialized (non-basic-OCR) Document AI processor. Model your actual document mix's split between plain-text and structured-extraction pages before choosing a vendor, since the cheaper option depends entirely on that ratio, not on a single vendor being universally cheaper.