OpenAI Prompt Caching vs Anthropic Prompt Caching vs Gemini Context Caching: Real Cost per 1 Billion Reused Tokens in 2026
The short answer: all three caching mechanisms look similar from a distance — reuse a prefix, pay less — but they bill the write step differently enough that the break-even point varies by vendor. OpenAI's current flagship (GPT-5.6 Sol) and Anthropic's current flagship (Claude Sonnet 5.5) both use the same 1.25x-write / 0.1x-read structure on their explicit-caching tiers, so they land on identical percentage savings: at a 10-reads-per-write reuse ratio, 1 billion reused tokens costs about $900 on GPT-5.6 Sol or $450 on Claude Sonnet 5.5 (the dollar gap reflects GPT-5.6 Sol's higher base rate, $4/M versus $2/M, not a different discount structure), against $4,400 and $2,200 respectively with no caching at all — roughly 80% savings on both. Gemini's current flagship (3.8 Flash Standard) charges no write premium at all (1.0x, the same as a normal input token) and, notably, now discounts reads by exactly the same 90% as Anthropic and OpenAI — a change from Gemini's older, shallower discount structure — but adds a separate, ongoing storage fee for keeping the cache alive that the other two don't carry. A workload with a low reuse ratio can still make caching a net cost on Anthropic's longer-TTL tier specifically, where the write premium is steeper.
What each vendor actually bills
OpenAI GPT-5.6 Sol (OFFICIAL, platform.openai.com/docs/pricing and platform.openai.com/docs/guides/prompt-caching). Input: $4/million tokens. Output: $20/million. Explicit caching (the current structure for GPT-5.6 Sol and newer): cache writes cost 1.25x the base input rate ($5/million), cache reads cost 0.1x ($0.40/million) — a 90% discount. An older, parallel automatic mechanism still runs on several non-flagship models: no separate write charge, and a read discount described as "up to 90%" that varies by model rather than a flat rate. Cached tokens apply to input only, and both cached and uncached tokens count toward account-level rate limits.
Anthropic Claude Sonnet 5.5 (OFFICIAL, docs.claude.com/en/docs/build-with-claude/prompt-caching). Input: $2/million tokens. Output: $10/million. Explicit, breakpoint-based caching (up to four cache boundaries per request): cache write at the default 5-minute TTL costs 1.25x base input ($2.50/million); an extended 1-hour TTL costs 2.0x ($4.00/million). Cache reads cost 0.1x base input ($0.20/million) — a 90% discount, identical in structure to OpenAI's current explicit-caching tier. Available identically through Anthropic's own API, AWS Bedrock, and Google Vertex AI at the same multipliers. Caching applies to input tokens only; cached tokens still count toward rate limits.
Google Gemini 3.8 Flash Standard (OFFICIAL, Google's current Gemini API pricing, introductory rate through December 31, 2026, after which it is scheduled to double). Input: $0.75/million tokens. Output: $3.75/million. Two caching mechanisms coexist. Implicit (automatic) caching requires no API changes and applies a discount automatically on a prefix match, at no separate storage charge, with a minimum cacheable prefix that varies by model (historically 1,024–4,096 tokens on recent Gemini generations). Explicit caching (the CachedContent API) gives a developer-controlled cache object with its own TTL: cache creation bills at the standard input rate with no write premium ($0.75/million, confirmed directly in Google's own caching documentation, which describes cache-token billing as applying "when included in subsequent prompts" rather than carrying its own separate creation surcharge); cache reads cost $0.075/million — a 90% discount, the same depth as Anthropic's and OpenAI's current rate, a change from Gemini's older generations which discounted reads more shallowly; and a separate, ongoing storage fee of $0.50 per million tokens per hour applies for every hour the cache object stays alive, independent of whether it's actually read during that hour. This storage fee is the one genuinely different cost mechanic in this comparison: Anthropic and OpenAI charge nothing to hold a cache that isn't being read; Gemini's explicit cache keeps billing until it's deleted or its TTL expires.
Normalizing the workload: reused tokens, write cost, and hit ratio
All ILLUSTRATIVE. This article defines "1 billion reused tokens" as 1 billion tokens' worth of cache reads and separately accounts for the write tokens needed to establish those caches. The connecting variable is the reuse ratio: how many cache reads occur, on average, per cache write. Formula: write tokens = reused tokens ÷ reuse ratio; total cost = (write tokens ÷ 1,000,000 × write rate) + (reused tokens ÷ 1,000,000 × read rate).
Cost per 1 billion reused tokens, at two reuse ratios
| Vendor / mechanism | Base input rate | Write rate | Read rate | Cost at 10:1 ratio | Cost at 100:1 ratio |
|---|---|---|---|---|---|
| Gemini 3.8 Flash Standard, explicit (token cost only, excludes storage) | $0.75/M | $0.75/M (1.0x) | $0.075/M (0.1x) | $150 | $83 |
| Claude Sonnet 5.5, 5-minute TTL | $2.00/M | $2.50/M (1.25x) | $0.20/M (0.1x) | $450 | $225 |
| Claude Sonnet 5.5, 1-hour TTL | $2.00/M | $4.00/M (2.0x) | $0.20/M (0.1x) | $600 | $240 |
| GPT-5.6 Sol, explicit cache | $4.00/M | $5.00/M (1.25x) | $0.40/M (0.1x) | $900 | $450 |
Gemini's lower total at this volume reflects its much lower base per-token rate for this model tier, not a deeper caching discount — its read discount (90%) now matches Anthropic's and OpenAI's exactly, a genuine convergence across all three vendors' current flagship-adjacent pricing.
Savings versus no caching at all
Formula: uncached cost = (write tokens + reused tokens) ÷ 1,000,000 × base input rate; savings % = (uncached − cached) ÷ uncached.
| Vendor / mechanism | Uncached cost, 10:1 ratio | Cached cost | Savings |
|---|---|---|---|
| Claude Sonnet 5.5, 5-minute TTL | $2,200 | $450 | 79.5% |
| GPT-5.6 Sol, explicit cache | $4,400 | $900 | 79.5% |
| Gemini 3.8 Flash Standard, explicit (excl. storage) | $825 | $150 | 81.8% |
At a 10:1 reuse ratio, Claude Sonnet 5.5's and GPT-5.6 Sol's identical 1.25x/0.1x structure produces the identical 79.5% savings rate — the percentage saved is structural, following directly from the multiplier pair, not vendor-specific. Savings improve further at the 100:1 ratio (88.9% on both), since the one-time write cost amortizes over more reads.
The exact break-even point, not a vague "several reads"
Formula: break-even reuse (R) = (write multiplier − 1) ÷ (1 − read multiplier) — the number of reads after which a cache write has fully paid for itself relative to never caching at all.
For Anthropic's and OpenAI's shared 1.25x-write / 0.1x-read structure: R = (1.25 − 1) ÷ (1 − 0.1) = 0.28, meaning a single subsequent read already pays back the 25% write premium — verified directly: write + 1 read = 1.25 + 0.1 = 1.35, against 2 uncached reads' worth (1.35 < 2). For Anthropic's 2.0x-write / 0.1x-read 1-hour TTL tier: R = (2.0 − 1) ÷ (1 − 0.1) = 1.11, meaning exactly two reads are required to beat the uncached cost — verified: write + 2 reads = 2.0 + 0.2 = 2.2, against 3 uncached reads' worth (2.2 < 3), while at one read it does not yet break even (2.0 + 0.1 = 2.1 against 2.0 uncached). Gemini's 1.0x write multiplier means it never costs more than uncached from a pure token-cost perspective — the break-even is zero reads — but its ongoing storage fee introduces a separate cost Anthropic's and OpenAI's write-premium model doesn't carry.
A realistic Gemini cache-lifetime example, storage included
Token cost alone understates Gemini's explicit-cache invoice. Formula: total cost = write (creation) + (reads × read rate) + (storage rate × hours held). For a 10,000-token system prompt cached for one hour and read 50 times during that hour: write = $0.0075, reads = 50 × 10,000 ÷ 1,000,000 × $0.075 = $0.0375, storage = 10,000 ÷ 1,000,000 × $0.50 × 1 hour = $0.005 — a total of $0.05 for that one cache's full one-hour lifetime, of which storage is 10% of the bill even in this read-heavy example. A cache held open for many hours with sparse reads shifts that proportion sharply toward storage, which is the genuinely distinguishing cost risk Gemini's explicit caching carries that Anthropic's and OpenAI's models do not.
The minimum cacheable prefix
Every vendor requires a minimum prefix length before caching activates — OpenAI's automatic mechanism needs at least 1,024 tokens, Anthropic's current models have historically required a comparable minimum, and Gemini's minimum has varied from 1,024 to 4,096 tokens depending on model generation. A short, frequently-changing system prompt under this threshold cannot be cached on any of these platforms regardless of how often it repeats.
Sensitivity
- Reuse ratio. The single largest lever in this comparison; low reuse narrows savings sharply, and on Anthropic's 1-hour TTL tier specifically, reuse below the two-read break-even turns caching into a net cost.
- Which OpenAI model generation is in use. The legacy automatic mechanism (no write charge) and the GPT-5.6-and-newer explicit mechanism (1.25x write charge) behave differently enough that model choice alone changes which pricing logic applies.
- Gemini's cache lifetime versus read frequency. A cache read frequently relative to its TTL amortizes the storage fee well; one left open with sparse reads accrues storage cost disproportionate to the savings it delivers.
- TTL choice on Anthropic. Moving from the 5-minute to the 1-hour TTL roughly doubles the write cost and raises the break-even from one read to two.
Budgeting traps
- Assuming caching always saves money regardless of reuse pattern. At a reuse ratio below the relevant break-even — exactly one read for the 1.25x tier, two for the 2.0x tier — the write premium can exceed what the discounted reads save.
- Ignoring Gemini's storage fee because the per-token read discount looks comparable to Anthropic's or OpenAI's. It is the one cost in this comparison that accrues independent of usage, for as long as the cache object exists.
- Treating OpenAI's caching as one single mechanism. The legacy automatic, no-write-fee behavior and the GPT-5.6-and-newer explicit 1.25x/0.1x structure are genuinely different cost models depending on which model is called.
- Forgetting that cached tokens still count toward rate limits. Caching reduces the bill, not the account's tokens-per-minute ceiling, on both OpenAI and Anthropic.
- Budgeting Gemini 3.8 Flash's introductory rate past its stated expiration. The $0.75/$3.75 rates are scheduled to double on January 1, 2027.
What to ask before you buy
Measure your own workload's actual reuse ratio (reads per write) on a representative sample before assuming caching will save a specific percentage, and compare it directly against the exact break-even point for your chosen TTL tier rather than assuming any reuse automatically pays off. If evaluating Gemini's explicit caching, confirm the current storage fee and your expected cache idle time, since that cost does not appear in Anthropic's or OpenAI's per-token-only pricing model.