Cheapest AI API in 2026: Real Cost per 1M Tokens Compared
If you only want the number: DeepSeek V4.1 Flash's off-peak rate — $0.15/M input, $0.60/M output — is the cheapest widely-available frontier-adjacent API on raw per-token pricing as of September 2026, narrowly ahead of GPT-5.6 Luna ($0.20/M input, $1.20/M output). Qwen-Turbo undercuts both, but sits in a lighter, different capability class not really comparable to general-purpose generation. Everything past that headline depends on your workload, your region's traffic pattern relative to time-of-day pricing, and whether "cheapest per token" actually means "cheapest per task" for what you're building — which it often doesn't.
This is a market-wide view. If you specifically need GPT-5.6 Luna vs. Claude Sonnet 5 vs. Gemini 3.7 Flash worked out across multiple real workloads, NotCheapAI has a dedicated comparison for that — this piece is the broader map: who else is on the board, what they actually cost, and where the real per-token floor sits right now.
The current market, side by side
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| DeepSeek V4.1 Flash (off-peak) | $0.15 | $0.60 | 1M | Cache hit: $0.003/M. Off-peak = 79% of the week |
| DeepSeek V4.1 Flash (peak) | $0.30 | $1.20 | 1M | Peak: 01:00–04:00 & 06:00–10:00 UTC, weekdays |
| GPT-5.6 Luna | $0.20 | $1.20 | 1.1M | Permanent rate since July 30, 2026 cut |
| Qwen-Turbo | ~$0.05 | ~$0.20 | Varies by tier | Alibaba's cheapest tier; classification/routing-grade |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1M | Promotional through Dec 31, 2026; doubles to $1.50/$7.50 after |
| GPT-5.6 Terra | $2.00 | $12.00 | 1.05M | Permanent rate since July 30, 2026 cut |
| Claude Sonnet 5 | $3.00 | $15.00 | 1M | Standard rate since Sept 1, 2026 (was $2/$10 intro through Aug 31) |
| Qwen3-Max | ~$1.20 | ~$6.00 | 262K | Alibaba's flagship-adjacent tier; check current Qwen guide for full tier ladder |
| GPT-5.6 Sol | $4.00 | $20.00 | 1.05M | Promotional through at least Nov 21, 2026; standard $5/$30 |
A note on Qwen: Alibaba runs one of the widest and fastest-moving tier ladders in the market — from a sub-$0.10 Turbo tier meant for classification and routing up through several Max-class variants that have shipped multiple times in 2026 alone. The pricing changes often enough, across enough tiers, that folding a full breakdown into this hub would go stale within weeks. NotCheapAI's dedicated Qwen pricing guide tracks the full tier ladder and regional pricing splits; the figures above are a current-as-of-writing snapshot for market context, not the full picture.
Why the "cheapest" answer changes with your workload
Every number in that table is a per-token rate. Your bill is per-token rate × your actual token volume, split across input and output — and that split varies enormously by task type. A model that wins on input pricing can lose badly on a job that's mostly output, and vice versa.
CALCULATION — three models, same workload shape (1M input tokens, 1M output tokens, a 1:1 blend used purely for comparison, not a claim about real traffic):
| Model | Input cost | Output cost | Blended total |
|---|---|---|---|
| DeepSeek V4.1 Flash (off-peak) | $0.15 | $0.60 | $0.75 |
| GPT-5.6 Luna | $0.20 | $1.20 | $1.40 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $4.50 |
| Claude Sonnet 5 | $3.00 | $15.00 | $18.00 |
At this 1:1 blend, DeepSeek V4.1 Flash off-peak is the cheapest, narrowly ahead of Luna. But shift the mix to something input-heavy — a long-document, short-answer task, 10M input tokens to 500K output tokens — and the gap between DeepSeek and Luna widens because DeepSeek's input rate is proportionally further ahead of Luna's than its output rate is:
- DeepSeek V4.1 Flash off-peak: (10 × $0.15) + (0.5 × $0.60) = $1.50 + $0.30 = $1.80
- GPT-5.6 Luna: (10 × $0.20) + (0.5 × $1.20) = $2.00 + $0.60 = $2.60
Flip it to output-heavy — 500K input, 10M output, a long-form generation task — and Luna's output rate (exactly 2x DeepSeek's off-peak output rate) starts to matter more:
- DeepSeek V4.1 Flash off-peak: (0.5 × $0.15) + (10 × $0.60) = $0.075 + $6.00 = $6.08
- GPT-5.6 Luna: (0.5 × $0.20) + (10 × $1.20) = $0.10 + $12.00 = $12.10
DeepSeek stays ahead in both directions here, but the margin changes a lot — roughly 31% cheaper in the input-heavy case, exactly 50% cheaper in the output-heavy case. That gap is the difference between "worth switching for" and "rounding error," and it only shows up when you model your actual token split instead of quoting the headline rate.
The DeepSeek caveat nobody's headline mentions
DeepSeek V4.1 Flash's $0.15/$0.60 rate is an off-peak rate. During peak hours — 01:00–04:00 UTC and 06:00–10:00 UTC, Monday through Friday, which lines up with Chinese business hours — both meters double, to $0.30/$1.20. That's still competitive with GPT-5.6 Luna at peak, but the cheapest-in-market claim specifically depends on either scheduling your traffic into the 79% of the week that's off-peak, or accepting that your blended real-world rate sits somewhere between the off-peak and peak numbers depending on your traffic's actual clock. NotCheapAI's dedicated DeepSeek V4.1 Flash pricing guide works through the full peak/off-peak mechanics, caching, and retry-cost math if that's the deciding factor for your use case.
The promotional-pricing trap
Three of the nine prices in the table above are explicitly time-limited promotions, and at least one widely-circulated comparison online is already citing an expired promotional rate as current:
- Claude Sonnet 5's $2/$10 rate expired August 31, 2026. Standard pricing of $3/$15 has applied since September 1, 2026 — that's two weeks ago as of this writing, and plenty of pricing content hasn't caught up.
- Gemini 3.7 Flash's $0.75/$3.75 rate is promotional through December 31, 2026, reverting to $1.50/$7.50 on January 1, 2027.
- GPT-5.6 Sol's $4/$20 rate is promotional, confirmed through at least November 21, 2026, with standard pricing at $5/$30.
If you're budgeting past those reversion dates, use the standard rate, not the promotional one — and if you're comparing "cheapest API" content published before September 2026, assume the Sonnet 5 figure in it is stale.
Realistic workload comparison across five models
Take a moderate content-operations workload: 2,000 requests/month, each with 3,000 input tokens and 2,000 output tokens (5,000 total tokens per request).
- Total input: 2,000 × 3,000 = 6,000,000 (6M) tokens
- Total output: 2,000 × 2,000 = 4,000,000 (4M) tokens
| Model | Input cost | Output cost | Total/month |
|---|---|---|---|
| DeepSeek V4.1 Flash (off-peak) | 6 × $0.15 = $0.90 | 4 × $0.60 = $2.40 | $3.30 |
| GPT-5.6 Luna | 6 × $0.20 = $1.20 | 4 × $1.20 = $4.80 | $6.00 |
| Gemini 3.7 Flash | 6 × $0.75 = $4.50 | 4 × $3.75 = $15.00 | $19.50 |
| GPT-5.6 Terra | 6 × $2.00 = $12.00 | 4 × $12.00 = $48.00 | $60.00 |
| Claude Sonnet 5 | 6 × $3.00 = $18.00 | 4 × $15.00 = $60.00 | $78.00 |
The spread from cheapest to most expensive here is roughly 24x, on the same exact workload, purely from model choice. That's the number worth internalizing: model selection is very often a bigger lever on cost than any prompt optimization, caching setup, or batch discount you could layer on top of a single model.
What the sticker price doesn't include
Caching. Every major provider offers some form of prompt or context caching, and the discount is substantial where it applies — DeepSeek's cache-hit rate is roughly 2% of its cache-miss rate; Gemini 3.7 Flash's cache read runs about 10% of base input. None of the totals above assume any caching, because caching benefit depends entirely on how much of your input is genuinely reused across calls — an assumption this hub can't make for you.
Batch discounts. Non-real-time workloads (evaluation runs, bulk processing, anything that doesn't need an immediate response) typically qualify for a flat discount, often around 50%, on providers that offer it. If your workload is schedulable, check for this before assuming the real-time rate is your floor.
Retries. A generation that fails validation, times out, or gets discarded still consumed billed tokens. If your pipeline has a meaningful retry rate, your effective cost is your table rate multiplied by roughly (1 + retry rate) — a 20% retry rate adds roughly 20% to your bill, concentrated on whichever meter (input or output) your retries resend.
Tool calls, if your use case needs them. Web search, code execution, and similar tool integrations are billed separately from token usage on every major API — commonly in the $2.50–$10 per 1,000 calls range depending on the tool and provider — and that's on top of the token cost for processing whatever the tool returns.
Where self-hosting open weights fits in
Every model in the comparison table above is a hosted API — someone else's infrastructure, billed per token. Several of the cheapest options (DeepSeek's underlying weights, Qwen's smaller tiers) are also available as open-weight downloads, which raises the question of whether self-hosting beats even the cheapest API rate.
The honest answer is that self-hosting isn't a per-token cost at all — it's a fixed infrastructure cost (GPU rental or ownership, at minimum several hundred dollars a month for hardware capable of running a large MoE model at usable throughput) plus the engineering time to deploy, monitor, and keep it running. Below a certain volume, that fixed cost is far more expensive than paying even the priciest API in this comparison. Above a very high, sustained volume — the kind where a fixed monthly GPU cost is dwarfed by the token volume it processes — self-hosting can undercut any hosted API on a per-token basis, because you're no longer paying anyone else's margin. Where exactly that crossover point sits depends entirely on your actual sustained volume and the GPU pricing available to you, and it's a large enough calculation that it deserves its own dedicated breakdown rather than a paragraph here — treat this section as a signpost, not a verdict: for the volume most readers of this guide are running, a hosted API at DeepSeek or Luna's rates is very likely still cheaper than standing up and maintaining your own inference infrastructure.
A word on rate volatility in this market
Every price in the comparison table above carries an implicit "as of September 15, 2026" timestamp, and that timestamp matters more in this market than in almost any other software category NotCheapAI covers. In the past few months alone: GitHub-adjacent coding tools have cut and raised prices within the same quarter, DeepSeek moved from flat pricing to a peak/off-peak structure and then raised both tiers substantially, OpenAI cut GPT-5.6 Luna's price by 80% five weeks after launch, and Anthropic's Sonnet 5 introductory rate expired on schedule while multiple third-party trackers kept citing it as current weeks later. None of that is unusual for this market — it's the norm. The practical takeaway: treat every number in this piece, and in any AI pricing comparison, as a snapshot rather than a fixed fact, and re-verify against the provider's own pricing page before committing meaningful volume or negotiating a contract around a specific rate.
Where the real floor is right now
For pure text generation with a schedulable, cache-friendly, non-latency-critical workload, DeepSeek V4.1 Flash's off-peak rate is the lowest verifiable floor in this comparison. For workloads that need to run at consistent low latency around the clock — where DeepSeek's peak-hour doubling becomes a real factor — GPT-5.6 Luna's flat rate, with no time-of-day variation, is the more predictable cheap option. Below both of those, ultra-lightweight tasks like classification and routing can go meaningfully cheaper still on tiers like Qwen-Turbo, though that's a different capability class entirely and not a fair substitute for general-purpose generation.
Who should pick which tier
High-volume, cost-sensitive, schedulable workloads: DeepSeek V4.1 Flash, scheduled off-peak where possible.
High-volume workloads that can't tolerate peak-hour price doubling: GPT-5.6 Luna, for its flat rate.
Classification, routing, tagging, and other lightweight structured tasks: the cheapest tier available from any provider that offers one — Qwen-Turbo class pricing is built for exactly this and is wasteful to pay frontier-model rates for.
Tasks where output quality directly affects downstream cost (fewer retries, less human correction, better first-pass results): a pricier mid-tier model, verified against your own task rather than assumed from reputation — see the effective-cost-per-successful-task discussion in NotCheapAI's GPT-5.6 Luna vs. Claude Sonnet 5 vs. Gemini 3.7 Flash comparison for the full mechanics of that trade-off.
FAQ
What's the single cheapest AI API right now? On raw per-token pricing, DeepSeek V4.1 Flash's off-peak rate ($0.15/M input, $0.60/M output), narrowly ahead of GPT-5.6 Luna ($0.20/M input, $1.20/M output) on a balanced workload, and by a wider margin on input-heavy workloads specifically.
Is Claude Sonnet 5 still $2/$10 per million tokens? No. That was an introductory rate that expired August 31, 2026. Standard pricing of $3/$15 has applied since September 1, 2026.
Why isn't Qwen ranked in detail in this comparison? Alibaba's Qwen lineup spans a very wide tier ladder that updates frequently — folding a precise, current breakdown into a broad market hub risks going stale fast. See NotCheapAI's dedicated Qwen pricing guide for the full current tier ladder.
Does the cheapest model per token always produce the cheapest result? No — see the effective-cost-per-successful-task principle above. A pricier model that needs fewer retries or produces more directly usable output can beat a cheaper model that needs several attempts. This is task-dependent and worth testing on your own workload rather than assuming.
How often does this kind of pricing comparison go stale? Judging by the number of promotional rates active in this market right now — three of the nine models compared here carry an explicit reversion date — expect meaningful movement within months, not years. Always check the provider's current pricing page before committing volume.
The NotCheapAI verdict
The cheapest AI API in 2026 isn't a single answer — it's DeepSeek V4.1 Flash if your traffic can work with its peak/off-peak schedule, GPT-5.6 Luna if it can't, and a much cheaper specialized tier still if your task doesn't need frontier-model capability at all. What's consistent across every workload modeled here is the size of the gap: choosing the right model class for the job is worth far more than any single optimization layered on top of the wrong one.