NotCheapMAKE EVERY CREDIT COUNT

DeepSeek V4.1 Flash Pricing: How Cheap Is the API Really in 2026?

DeepSeek V4.1 Flash costs $0.15 per million input tokens and $0.60 per million output tokens during off-peak hours. During peak hours — 01:00–04:00 UTC and 06:00–10:00 UTC, Monday through Friday — both rates double, to $0.30 and $1.20. Cache-hit input tokens cost a fraction of a cent: $0.003 off-peak, $0.006 peak.

Those four numbers are the whole headline, and on their own they make DeepSeek V4.1 Flash one of the cheapest frontier-adjacent APIs on the market. The real question — the one that actually determines your invoice — is what happens once you add real traffic patterns, cache-miss rates, and the retries every production system generates. That's what this guide works through.

The full DeepSeek V4.1 Flash pricing table

Meter (per 1M tokens)Off-peakPeak
Input, cache hit$0.003$0.006
Input, cache miss (uncached)$0.15$0.30
Output$0.60$1.20
Context window1,000,000 tokens1,000,000 tokens
Max output per request384,000 tokens384,000 tokens
Concurrency limit2,5002,500

FACT: Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC, Monday to Friday — seven hours a day, five days a week. Every other hour, including all of Saturday and Sunday, is off-peak.

CALCULATION: That's 35 peak hours out of a 168-hour week — roughly 21% of the week is peak, and about 79% is off-peak. If your traffic isn't pinned to a clock, the odds already favor the cheaper rate.

ASSUMPTION worth flagging: the peak window is set in UTC and lines up with Chinese business hours (roughly 09:00–12:00 and 14:00–18:00 in Beijing). For a US-based product, both windows fall in the middle of the US night, so daytime American traffic lands almost entirely in off-peak. For a UK or continental European product, the second window (06:00–10:00 UTC) overlaps with the start of the European workday — so a London-based team running live daytime traffic will actually eat some peak pricing, unlike a US-based one running the same hours.

Since September 10, 2026, this pricing applies under the model name deepseek-flash, replacing the previous V4 Flash line outright. From noon Beijing time on September 14, 2026, DeepSeek also began routing all deepseek-v4-pro requests to V4.1 Flash, billed at V4.1 Flash's rates, until a V4.1 Pro model is released. DeepSeek has not given a date for that. If you're still calling the old Pro model string, you are — as of this writing — paying Flash prices for what used to be Pro-tier requests. That's a temporary state, not a permanent price cut, and it's worth watching for when V4.1 Pro arrives with its own (likely higher) price card.

Uncached input vs. cache-hit input: where the real savings live

The headline $0.15 figure only applies to tokens DeepSeek hasn't seen before in your current context — a cache miss. Once a chunk of input has been processed and is reused (a system prompt, a long document you're asking multiple questions about, a repeated conversation history), it becomes a cache hit and drops to $0.003 off-peak: 98% cheaper than the uncached rate.

Scenario (off-peak)Cost per 1M input tokens
100% cache miss (first call, fresh input)$0.15
100% cache hit (fully reused context)$0.003
80% hit / 20% miss (typical repeated-context app)$0.033
50% hit / 50% miss$0.077

CALCULATION for the 80/20 blend: 0.8 × $0.003 + 0.2 × $0.15 = $0.0024 + $0.03 = $0.0324 per million tokens, rounded to $0.033 above.

Caching helps most in exactly one pattern: a large, stable piece of context (a knowledge base excerpt, a codebase, a long system prompt) that gets reused across many calls with only a small amount of new text appended each time. It helps least when every request is genuinely novel — one-off summarization of unique documents, for example, where there's nothing to hit a cache with. Don't assume caching will bail out an expensive workload; check whether your actual traffic has repeated context to hit in the first place.

Why output tokens usually dominate the bill

Output is priced at 4x the uncached input rate off-peak ($0.60 vs. $0.15) and stays at that 4x ratio at peak ($1.20 vs. $0.30). For most real applications, output is also where the token volume concentrates — a chatbot reply, a generated article, a piece of code, a long-form answer.

CALCULATION — a request with 2,000 input tokens (cache miss) and 2,000 output tokens, off-peak:

  • Input: 2,000 × ($0.15 / 1,000,000) = $0.0003
  • Output: 2,000 × ($0.60 / 1,000,000) = $0.0012
  • Output is 80% of this request's cost, despite equal token counts.

The imbalance grows further with any request where output is proportionally larger than input — long-form generation, code completion, or detailed explanations from a short prompt. If you're trying to control spend, the highest-leverage move is usually capping or trimming output length and being deliberate about max_tokens, not shaving input.

Peak vs. off-peak economics

Monthly workloadOff-peak costPeak costDifference
Low (1M input / 0.5M output, cache miss)$0.45$0.902x
Medium (10M input / 5M output, cache miss)$4.50$9.002x
High (100M input / 50M output, cache miss)$45.00$90.002x

CALCULATION for "Medium": off-peak = (10 × $0.15) + (5 × $0.60) = $1.50 + $3.00 = $4.50. Peak = (10 × $0.30) + (5 × $1.20) = $3.00 + $6.00 = $9.00.

Because peak is a flat doubling of every meter, the math scales identically at every volume — there's no tier where peak pricing becomes proportionally worse or better. The only variable that matters is what fraction of your actual traffic falls inside those seven daily hours.

A realistic monthly workload, worked in full

Take a small content-operations use case: a team generating 400 draft articles a month, where each generation call sends roughly 3,000 tokens of input (a brief, style guide, and reference material — much of it reused across calls) and returns roughly 3,500 tokens of output per article.

  • Total input: 400 × 3,000 = 1,200,000 tokens
  • Total output: 400 × 3,500 = 1,400,000 tokens

Scenario A — no caching, all off-peak:

  • Input: 1.2 × $0.15 = $0.18
  • Output: 1.4 × $0.60 = $0.84
  • Total: $1.02/month

Scenario B — 70% of input is reused style-guide/reference material (cache hit), all off-peak:

  • Cache-hit input: 0.84M × $0.003/1M = $0.0025
  • Cache-miss input: 0.36M × $0.15/1M = $0.054
  • Output: 1.4 × $0.60 = $0.84
  • Total: ≈ $0.90/month

Scenario C — same as B, but 20% of generations need a retry:

  • Effective output: 1.4M × 1.2 = 1.68M → 1.68 × $0.60 = $1.008
  • Effective cache-miss input: 0.36M × 1.2 = 0.432M → 0.432 × $0.15 = $0.065
  • Effective cache-hit input: 0.84M × 1.2 = 1.008M × $0.003/1M = $0.003
  • Total: ≈ $1.08/month

Even with a 20% retry rate stacked on top, this specific workload stays under $1.10 a month — which is the honest picture of why DeepSeek V4.1 Flash gets described as "cheap": at this kind of volume, the absolute numbers are small enough that even a meaningfully inefficient pipeline barely registers. The picture changes at 100x this volume, where the same 20% retry overhead becomes a real, budgeted line item rather than rounding error — which is exactly why the retry-rate math earlier in this guide matters more as volume scales up.

What batch and schedulable workloads can save

If your workload doesn't need to run in real time — nightly data processing, bulk content generation, embedding a document archive, running eval suites — you can pin it entirely to off-peak hours and guarantee the cheaper rate rather than hoping your traffic happens to land there.

CALCULATION: a 500M-token batch job (400M input cache-miss, 100M output) run entirely off-peak costs (400 × $0.15) + (100 × $0.60) = $60 + $60 = $120. The same job run entirely at peak costs (400 × $0.30) + (100 × $1.20) = $120 + $120 = $240. Scheduling the job into the 133 off-peak hours of the week — which is most of it — saves $120 outright, for zero change to the work itself.

This is the clearest "free money" available in DeepSeek's pricing model: if a job can wait a few hours, there's rarely a reason to run it at peak.

Retries and failed generations: the hidden line item

DeepSeek bills for tokens generated, not tokens used. A response that fails validation, gets rejected by a downstream check, times out, or simply isn't what you needed still consumed real output tokens — and often real input tokens too, if the retry re-sends context.

CALCULATION: take the "Medium" workload above ($4.50/month off-peak at a 0% retry rate) and add a 20% retry rate, where a fifth of all generations have to be redone:

  • Effective output tokens billed: 5M × 1.20 = 6M
  • Effective input tokens billed (assuming retries resend the same input): 10M × 1.20 = 12M
  • New cost: (12 × $0.15) + (6 × $0.60) = $1.80 + $3.60 = $5.40, a 20% increase that tracks the retry rate almost exactly.

Retry rate is invisible in a pricing page but very visible on an invoice. If you're evaluating DeepSeek V4.1 Flash against a pricier alternative, the honest comparison isn't sticker price vs. sticker price — it's sticker price adjusted for how often each model needs a second attempt to get a usable result. That figure isn't published by any vendor and has to come from your own testing.

How DeepSeek V4.1 Flash compares to other current APIs

ModelInput $/MOutput $/MNotes
DeepSeek V4.1 Flash (off-peak)$0.15$0.60Cache hit drops input to $0.003
DeepSeek V4.1 Flash (peak)$0.30$1.207 hrs/day, weekdays only
GPT-5.6 Luna$0.20$1.20OpenAI's high-volume, cost-tier model
Gemini 3.7 Flash$0.75$3.75Promotional rate through Dec 31, 2026

At off-peak rates, DeepSeek V4.1 Flash undercuts GPT-5.6 Luna on input and matches it on output, and it undercuts Gemini 3.7 Flash on both meters by a wide margin. At peak rates, it's still cheaper than Gemini 3.7 Flash on both meters and roughly comparable to GPT-5.6 Luna, slightly more expensive on input and matching on output. The gap narrows considerably once you're paying peak prices, which is exactly why scheduling matters so much for this specific model.

Where DeepSeek V4.1 Flash looks exceptionally cheap

  • Batch and offline workloads you can pin to off-peak hours — the savings are guaranteed, not probabilistic.
  • High-cache-reuse applications — a support bot answering questions against the same knowledge base, a coding assistant working repeatedly inside the same large codebase — where the $0.003 cache-hit rate does most of the heavy lifting.
  • High-volume, cost-sensitive workloads where the 1M-token context window lets you avoid chunking large documents into smaller, more expensive calls elsewhere.
  • Teams whose real traffic clock genuinely misses the peak window — US-based daytime products, in particular.

Where choosing on token price alone would be a mistake

  • Latency- or correctness-critical live traffic during the peak window, especially for European teams whose morning hours overlap the 06:00–10:00 UTC slot.
  • Workloads with high retry rates on complex tasks. If a pricier model needs meaningfully fewer attempts to get a usable answer, its effective cost per successful output can undercut a cheaper model that needs several tries — this is a real economic effect even though no verified head-to-head retry-rate benchmark exists publicly for V4.1 Flash against GPT-5.6 or Gemini at the time of writing.
  • Regulated or compliance-sensitive use cases where data handling and residency terms matter more than the per-token rate — that's a due-diligence question separate from cost, and worth resolving with your own legal review before committing volume.
  • Anyone still pointed at deepseek-v4-pro expecting Pro-tier behavior: you're currently getting Flash-tier model behavior at Flash-tier prices, which may or may not match what your application was tuned for.

Who should use DeepSeek V4.1 Flash

Cost-sensitive teams running high-volume text workloads — classification, summarization, chat, content drafting, coding assistance — especially where the workload is schedulable, cache-friendly, or simply large enough that a fraction-of-a-cent difference per million tokens compounds into real monthly savings.

Who should probably choose something else

Teams needing guaranteed low latency on live traffic that regularly falls inside the peak window, teams with compliance requirements the DeepSeek terms don't satisfy, and teams whose workloads have shown — through their own testing, not a marketing claim — meaningfully higher retry or correction rates on V4.1 Flash than on a pricier alternative for their specific task.

FAQ

Is DeepSeek V4.1 Flash the cheapest AI API available right now? On raw off-peak per-token pricing, it's among the cheapest frontier-adjacent options publicly available, undercutting both GPT-5.6 Luna and Gemini 3.7 Flash on at least one meter. Whether it's the cheapest for your specific workload depends on your cache-hit rate, your peak/off-peak split, and your retry rate — none of which show up in the sticker price.

What happens to my deepseek-v4-pro calls now? As of noon Beijing time on September 14, 2026, they're being served by V4.1 Flash and billed at V4.1 Flash rates. There's no announced date for a V4.1 Pro release.

Does the peak/off-peak split actually save money, or is it a marketing trick? It's a real, structural discount if your traffic can shift — 79% of the week is off-peak, and the off-peak rate is exactly half the peak rate on every meter. If your traffic can't shift (real-time, always-on, globally distributed), the split doesn't help you and you should budget for a peak-weighted average instead.

Is the $0.003 cache-hit rate realistic for most apps, or a best-case number? It's real, but it only applies to the portion of input that's actually a cache hit. An app with no repeated context structure won't see it at all; an app built around a large reused system prompt or knowledge base will see it dominate their input cost.

The NotCheapAI verdict

DeepSeek V4.1 Flash's headline pricing is genuinely cheap, and the peak/off-peak structure rewards exactly the kind of workload most cost-sensitive teams already run: batchable, cache-friendly, non-latency-critical text processing. But the sticker price is a starting point, not an invoice. The three multipliers that actually decide what you pay — your peak/off-peak traffic split, your cache-hit rate, and your retry rate — are all things you control or measure, not things DeepSeek publishes for you. Run the math on your own traffic pattern before assuming the cheapest number on the page is the cheapest number on your bill.