DeepSeek in 2026: How Cheap Is It Really?
DeepSeek's chat app is free to try, but its developer API is metered. The current low-cost API model is deepseek-flash, serving DeepSeek-V4.1-Flash. At the official off-peak price, one million uncached input tokens cost $0.15 and one million output tokens cost $0.60. During weekday peak hours, those rates double. The separate deepseek-v4-pro API model remains listed at higher prices. An old model name in working code is not evidence that the old model is still serving the request.
What is free, and what is billed?
The DeepSeek website links to its web assistant and API platform as separate products; its official app listing advertises free assistant use. This does not include a developer API allowance or promise that every chat feature has unlimited capacity. DeepSeek's API documentation says token charges come from a topped-up or granted balance. Pay-as-you-go billing begins with actual API calls, independent of whether you use the app.
The current model and rate card
All prices below are US dollars per one million tokens, from DeepSeek's Models & Pricing page. Input cache hit and cache miss are separate billing categories. Output has its own rate.
| API model | Time | Input: cache hit | Input: cache miss | Output |
|---|---|---|---|---|
deepseek-flash (V4.1-Flash) | Off-peak | $0.003 | $0.15 | $0.60 |
deepseek-flash (V4.1-Flash) | Peak | $0.006 | $0.30 | $1.20 |
deepseek-v4-pro (V4-Pro-0813) | Off-peak | $0.022 | $0.66 | $1.98 |
deepseek-v4-pro (V4-Pro-0813) | Peak | $0.044 | $1.32 | $3.96 |
Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday, excluding Chinese public holidays; all other hours are off-peak. A workload spanning both periods gets a blended bill. DeepSeek lists 1M context and a 384K maximum output for both API models, plus concurrency limits of 2,500 for Flash and 500 for Pro. These are service limits, not a guarantee that a maximum-size request is the economical way to use the model.
The model names matter. DeepSeek's September change log says the former V4 Flash and V4 Flash Vision Exp models are retired. Their legacy IDs may still resolve, but are routed to V4.1-Flash and billed at the Flash rate. Use deepseek-flash for a new integration. The same change log says the earlier proposal to stop V4 Pro service after September 14 was reversed: Pro remains available until further notice.
Cache hits are real savings, not a guaranteed discount
DeepSeek's context-caching guide says caching is automatic and based on matching input prefixes. It reports prompt_cache_hit_tokens and prompt_cache_miss_tokens in API usage. A cache hit is not guaranteed just because two prompts discuss the same topic: the prefix must match a persisted cache unit, and old units can be cleared after hours or days. Output tokens remain separately billed.
For a simple off-peak example, consider one million input tokens and one million output tokens accumulated over requests. If none of the input hits cache, Flash costs $0.15 + $0.60 = $0.75. If half the input actually hits cache, it costs 0.5 x $0.003 + 0.5 x $0.15 + $0.60 = $0.6765. The exact mix should come from response usage, not an assumed cache-hit percentage. Longer reasoning can also increase billed output volume, so compare completed tasks as well as token rates.
Worked API costs
The next estimates assume text input, no cache hits, no separate infrastructure, and fixed token counts. They show both clocks because off-peak is not a universal rate.
| Workload | Flash off-peak | Flash peak | Pro off-peak | Pro peak |
|---|---|---|---|---|
| One request: 10K input + 2K output | $0.00270 | $0.00540 | $0.01056 | $0.02112 |
| 1,000 such requests | $2.70 | $5.40 | $10.56 | $21.12 |
| Month: 400M input + 100M output | $120 | $240 | $462 | $924 |
For Flash's first cell: 0.01 x $0.15 input + 0.002 x $0.60 output = $0.00270. The monthly line multiplies rates by 400 and 100 respectively; it is a sum of many requests, not one impossible 400M-token prompt. Real bills may sit between the peak and off-peak columns. If your output is longer than assumed, cost rises quickly even when inputs are cheap.
What about hosting it yourself?
DeepSeek publishes V4.1-Flash weights under an MIT license. That can remove DeepSeek's per-token API charge for teams able to operate their own inference. It does not remove compute, memory, storage, staff time or reliability work. The published model card describes a 552B-parameter backbone, so self-hosting this specific model is not a single ordinary consumer-GPU project. Smaller or quantized checkpoints are a different capability and performance decision; do not compare their hardware bill with the official API as if the model were identical.
At the modeled 400M-input/100M-output month, official Flash costs $120 off-peak or $240 peak, before cache savings. That is a demanding benchmark for any always-on self-hosted stack to beat on dollars alone. Self-hosting can still make sense for control or deployment requirements, but its break-even point needs measured throughput and an actual hardware quote.
Is DeepSeek still cheap?
For metered text workloads, yes: Flash's off-peak API rate is low enough that even substantial usage can cost less than a consumer subscription elsewhere. But the conclusion has conditions: compare the right model, clock, cache behavior and output length. Pro is materially more expensive than Flash, and neither API price implies a quality guarantee for a particular task. The Qwen regional pricing guide shows another low-cost option; our four-provider comparison prices the same example work across providers.
Frequently asked questions
Is deepseek-v4-flash the current Flash model? No. The old ID may still work for compatibility, but DeepSeek says it now routes to V4.1-Flash. New code should request deepseek-flash.
Was V4 Pro retired in September? No. DeepSeek's current pricing page still lists it, and the change log says service continues until further notice.
Does the $0.003 cache-hit price apply to every Flash input token? Only input tokens reported as cache hits off-peak. Cache misses and all output tokens use their respective rates.