NotCheapMAKE EVERY CREDIT COUNT

OpenAI Batch API vs Anthropic Message Batches vs Google Gemini Batch: Real Batch LLM Cost per 1 Billion Tokens in 2026

The short answer: all three major providers discount asynchronous batch processing by exactly 50% off their standard synchronous rate, on both input and output tokens, with no exceptions found for any current-generation model from OpenAI, Anthropic, or Google. At an 80:20 input-to-output ratio, 1 billion tokens on a comparable mid-tier model (OpenAI's GPT-6 Sol or Anthropic's Claude Sonnet 5, both priced at $2/$10 per million tokens) costs $3,600 synchronous or $1,800 on batch; on Google's cheapest production-ready tier, Gemini 2.5 Flash-Lite, the same workload costs $160 synchronous or $80 on batch. The discount percentage is identical across vendors; the underlying model's own rate is what actually decides your bill.

What each vendor bills, and the discount mechanic

OpenAI Batch API (OFFICIAL). A flat 50% discount on both input and output tokens versus the synchronous rate for the same model, with results returned within a 24-hour completion window. OpenAI's own documentation describes this as "a 50% cost discount compared to synchronous APIs." Jobs are submitted as a file of requests rather than individual live calls; there is no separate per-job platform fee beyond the discounted token rate itself.

Anthropic Message Batches (OFFICIAL). Also a flat 50% discount on both input and output tokens, applying uniformly across Anthropic's current model lineup. One source reports Anthropic's batch turnaround as roughly 4 hours rather than OpenAI's 24-hour window, though this specific SLA figure should be confirmed against Anthropic's own current documentation before being relied on for a time-sensitive pipeline. Anthropic additionally supports combining batch pricing with prompt caching, where a cached-and-batched request can cost as little as roughly 5% of a standard non-cached, non-batched request on the same model — a stacking discount neither OpenAI nor Google was confirmed to offer at the same depth in the sources reviewed.

Google Gemini Batch API (OFFICIAL). Google's own pricing page describes its batch lane as delivering a "50% cost reduction" across Gemini models, matching the other two vendors exactly. One pricing tracker noted an additional, time-limited introductory rate on a specific Gemini Flash variant (reported at $0.375 per million input tokens through December 31, 2026, described as roughly 75% below the model's 2027 list price) that stacks on top of the standard batch discount — a temporary promotional rate, not part of Google's baseline batch mechanic, and one that should be reconfirmed against the live pricing page given how quickly Gemini's Flash-tier pricing has moved in 2026.

Reference model rates used in this comparison

All OFFICIAL or REPORTED as noted, current as of September 2026. OpenAI GPT-6 Sol Standard: $2.00 input / $10.00 output per million tokens synchronous; $1.00 / $5.00 on Batch. Anthropic Claude Sonnet 5 (REPORTED): $2.00 input / $10.00 output synchronous; $1.00 / $5.00 on Batch — coincidentally identical to OpenAI's mid-tier rate at the time of writing. Google Gemini 2.5 Flash-Lite: $0.10 input / $0.40 output synchronous; $0.05 / $0.20 on Batch, Google's cheapest production-ready tier and roughly 20 times cheaper per token than the OpenAI and Anthropic models used here — this is a genuine capability-tier difference, not a discount-structure difference, and this article makes no claim that the three models are interchangeable in quality.

Cost per volume, at an 80:20 input:output ratio

Formula: cost = (input millions × input rate) + (output millions × output rate), ILLUSTRATIVE 80% input / 20% output split.

Model100M tokens, sync100M tokens, batch1B tokens, sync1B tokens, batch10B tokens, sync10B tokens, batch
OpenAI GPT-6 Sol$360$180$3,600$1,800$36,000$18,000
Anthropic Claude Sonnet 5$360$180$3,600$1,800$36,000$18,000
Google Gemini 2.5 Flash-Lite$16$8$160$80$1,600$800

At every volume, the batch column is exactly half the sync column on every vendor, confirming the flat 50% mechanic holds across the full range tested. The gap between Gemini's cheap tier and the two mid-tier models here (22.5x at every volume) dwarfs the batch-versus-sync gap (2x); model selection is the bigger lever in this whole category, batch processing is the second-biggest.

The truncation trap

Batch pricing charges for the full generated response, not the truncated portion your application actually uses. A job that generates a 2,000-token response but only consumes the first 200 tokens after truncation is still billed for all 2,000 output tokens — a mechanic that applies identically across OpenAI, Anthropic, and Google, since batch pricing bills what the model generates, not what the caller keeps. Setting an explicit max_tokens or equivalent output-length cap before submission, rather than truncating after the fact, is the only way to avoid paying for output you discard.

Stacking batch with caching

On Anthropic specifically, batch pricing and prompt caching can combine: a cached-and-batched request was reported at as little as roughly 5% of the fully synchronous, non-cached rate on the same model — a saving well beyond the flat 50% batch discount alone. This matters most for workloads with a large, repeated static prefix (system prompts, long reference documents, few-shot examples) processed across many batch requests. OpenAI's automatic prompt-caching discount (reported up to 60% off repeated-prefix input tokens) and Google's explicit context-caching tier were both referenced as available, but a confirmed combined batch-plus-cache stacking rate comparable to Anthropic's ~5% figure was not found for either in the sources reviewed.

Sensitivity

  1. Input:output ratio. A workload skewed toward output (summarization-to-length tasks, code generation) pays proportionally more, since output tokens are priced 5x input tokens on both the OpenAI and Anthropic models used here.
  2. Which model tier is chosen. The choice between a frontier model and a cheap tier like Gemini Flash-Lite moves the bill by an order of magnitude or more, independent of batch versus sync.
  3. Cache hit rate. On Anthropic in particular, a high cache-hit workload can shrink the effective batch rate far below the flat 50% figure this article uses as its headline number.
  4. Turnaround tolerance. OpenAI's 24-hour window and Anthropic's reported ~4-hour window are not interchangeable for a pipeline with a hard latency requirement; confirm the current SLA for each vendor before committing a time-sensitive job to batch.

Budgeting traps

  • Forgetting batch bills the full generated output, not the truncated portion used. Set an explicit output-length limit before submission, not after.
  • Assuming every provider's batch discount is the same depth. All three match at a flat 50% baseline, but Anthropic's cache-plus-batch stacking can go considerably deeper; relying on that depth from another vendor without confirming it first will overstate expected savings.
  • Treating a promotional introductory rate as the standing batch price. The reported time-limited Gemini Flash rate expires December 31, 2026, per the source that flagged it; reconfirm before budgeting past that date.
  • Comparing batch rates across models of different capability tiers as if the discount alone determines value. Gemini Flash-Lite's headline advantage here comes overwhelmingly from its base rate, not from any batch-specific mechanic.

What to ask before you buy

Confirm each vendor's current batch turnaround window against your pipeline's latency tolerance, since a 24-hour and a 4-hour SLA are not interchangeable for anything with a same-day deadline. If your workload has a large repeated prefix, ask specifically about combined batch-plus-cache pricing, since that stacking discount — not the flat 50% batch rate alone — is where the largest additional savings in this category live.