Kimi K3 vs DeepSeek vs Qwen 2026: Real API Cost
Short Answer
The "cheap Chinese LLM" story no longer applies to all three labs equally. Kimi K3, Moonshot's new flagship released July 16, 2026, prices at $3.00 input / $15.00 output per million tokens — explicitly compared by outside observers to Anthropic's Sonnet tier, not to DeepSeek's or Qwen's budget rates. Its predecessor family (Kimi K2.6 / K2.7 Code) remains cheap at $0.95/$4.00. DeepSeek and Qwen's cheapest current tiers are still at the market floor — DeepSeek's Flash model runs $0.15–$0.30/M input depending on time of day, and Qwen's Plus tier runs about $0.40/M input on promotion — but each lab's own flagship model has climbed well past "cheap": DeepSeek V4 Pro reaches $0.66–$1.32/M input at peak, and Qwen3.8-Max sits at $2.00/M input, comparable to Kimi K3's input rate.
At the three normalized workloads in this brief:
| Workload | Kimi K3 | DeepSeek Flash (off-peak) | DeepSeek Pro (off-peak) | Qwen Plus (promo) | Qwen Max |
|---|---|---|---|---|---|
| 10M in + 2M out | $60.00 | $2.70 | $10.56 | $7.20 | $32.00 |
| 100M in + 20M out | $600.00 | $27.00 | $105.60 | $72.00 | $320.00 |
| 1B in + 200M out | $6,000.00 | $270.00 | $1,056.00 | $720.00 | $3,200.00 |
Kimi K3 is 20–22× DeepSeek's cheapest Flash tier at every volume, a gap that did not exist in the K2 generation. None of this is a quality judgment — it is a pricing-structure finding, and the structures themselves (three genuinely different billing mechanics) are the second half of this article.
1. Current Model Lineups and API Mechanics
Kimi / Moonshot — (platform.kimi.ai/docs/pricing/chat, fetched Sept 22, 2026)
| Model | Input (cache-miss) | Cached input | Output | Context |
|---|---|---|---|---|
| kimi-k3 | $3.00/M | $0.30/M (cache write: $3.00/M at 5-min TTL, $6.00/M at 1-hour TTL) | $15.00/M | 1,048,576 tokens |
| kimi-k2.7-code | $0.95/M | $0.19/M | $4.00/M | 262,144 tokens |
| kimi-k2.7-code-highspeed | $1.90/M | $0.38/M | $8.00/M | 262,144 tokens |
| kimi-k2.6 | $0.95/M | $0.16/M | $4.00/M | 262,144 tokens |
K3 has no non-thinking mode. reasoning_effort is adjustable (low/high/max, default max) but thinking cannot be disabled, and reasoning tokens bill as output at the full $15/M rate — a structural cost risk third-party reviewers flag repeatedly: a verbose reasoning trace can cost more than the visible answer. K3 is not currently listed for Batch pricing ( across multiple trackers; not stated as a limitation on the primary pricing page, which simply omits K3 from any batch table). K2.7 Code, K2.6 and the retired K2.5 do support Batch, priced at "60% of standard price" per secondary trackers (Moonshot's pricing page does not list K3 batch pricing).
Open weights ( across multiple independent trackers, consistent dates): K3's weights shipped July 27, 2026 under a Kimi K3 License (a modified-MIT-style license, not the plain MIT some Kimi materials reference elsewhere) — permitting use, modification, distribution and fine-tuning, but requiring a separate Moonshot agreement for Model-as-a-Service businesses earning over $20M in a trailing 12-month period, and mandatory attribution for very large commercial products. Self-hosting K3 at full precision reportedly requires substantial multi-GPU hardware. Published estimates range from 8 to 32 high-end accelerators depending on precision, so there is no single reliable configuration figure.
Reported rate limits (not listed on the cited official pricing page): a Tier 0–5 system gated by cumulative account recharge (roughly $1 to $3,000+), with concurrency, requests-per-minute and tokens-per-minute scaling by tier; a $1 minimum top-up is required before any request succeeds, and there is no free API tier.
Data residency remains unclear: one privacy-focused source says Moonshot stores API data in Singapore, while other sources describe it as China-hosted. The cited public materials do not settle the question; confirm Moonshot's current terms before using the service for regulated data.
DeepSeek — (api-docs.deepseek.com/quick_start/pricing, fetched Sept 22, 2026)
| Model | Input, cache-hit | Input, cache-miss | Output | Context |
|---|---|---|---|---|
| deepseek-flash (DeepSeek-V4.1-Flash) | off-peak $0.003 / peak $0.006 | off-peak $0.15 / peak $0.30 | off-peak $0.60 / peak $1.20 | 1M tokens, 384K max output |
| deepseek-v4-pro (DeepSeek-V4-Pro-0813) | off-peak $0.022 / peak $0.044 | off-peak $0.66 / peak $1.32 | off-peak $1.98 / peak $3.96 | 1M tokens, 384K max output |
Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday, excluding Chinese public holidays; every other hour — including weekends and full Chinese holidays — is off-peak at exactly half the peak rate. This is a genuine, official, published time-of-day billing structure, not a promotion. Concurrency limits: 2,500 for Flash, 500 for Pro. Both models support OpenAI-compatible, Anthropic-compatible and native Responses-API formats, Vision (Flash only), Tool Calls, JSON Output and Chat Prefix Completion. The legacy names deepseek-chat, deepseek-reasoner and deepseek-v4-flash are still accepted but now route to and bill as V4.1-Flash. DeepSeek states expenses are simply tokens × price and reminds developers that "product prices may vary" and to check the live page regularly.
Qwen / Alibaba Cloud Model Studio — (alibabacloud.com/help/en/model-studio/model-pricing, fetched Sept 22, 2026)
Qwen's pricing surface is an order of magnitude more complex than Kimi's or DeepSeek's: dozens of model-ID variants, five+ geographic regions (Singapore/International, China Beijing, Hong Kong, Germany/Frankfurt, US/Virginia, Japan/Tokyo) each with different prices for the same model, and input-length-tiered pricing on most models (the price per token rises once a single request crosses 32K, 128K or 256K input tokens — a structure neither Kimi nor DeepSeek uses).
Singapore/International rates (the region most non-China developers default to, and the one carrying a free quota):
| Model | Input (≤ first tier) | Output (≤ first tier) | Notes |
|---|---|---|---|
| qwen3.8-max | $2.00/M | $6.00/M | Flat across full 1M context (no tiering) |
| qwen3.7-max | $2.50/M | $7.50/M | Flat across full 1M context |
| qwen-max (legacy, non-thinking only) | $1.60/M | $6.40/M | No tiered pricing |
| qwen3.7-plus | $0.40/M (0–256K) → $1.20/M (256K–1M) | $1.60/M (0–256K) → $4.80/M (256K–1M) | Official page wording creates a discount/list-price ambiguity: it references a limited-time 20% discount, but this article does not treat the displayed $0.40 figure as settled beyond the captured Singapore/International table |
| qwen-plus (current alias) | $0.40/M (0–256K) → $1.20/M (256K–1M) | $1.20/M non-thinking, $4.00/M thinking (0–256K) | Thinking and non-thinking output billed differently — a structure unique to Qwen among the three labs |
Every price above is Singapore/International only. The same model IDs bill at different rates in China (Beijing), Hong Kong, Frankfurt, Virginia and Tokyo — for example, qwen3-max in China Beijing bills $0.359–$1.004/M input depending on request size, versus $1.20–$3.00/M in Singapore for the same tiers. Frankfurt, Hong Kong, Virginia and Tokyo additionally carry night/daytime discount pricing (night = 22:00–08:00 UTC+8, up to 60% off list) that Singapore does not offer. This article uses Singapore/International throughout for comparability with Kimi and DeepSeek's single global rate; do not assume these numbers apply to a China-region or EU-region deployment.
Batch calls: 50% of real-time price (vendor-listed, stated directly on the pricing page, applies where supported). Cached input: 10–20% of standard input price depending on model and cache type (explicit vs. implicit), with several current-generation models (qwen3.8-max, qwen3.8-flash, deepseek models hosted on the same platform) explicitly flagged as not following the standard 10–20% ratio. The exact cache rates for those models are account-specific and are not published as a single public number. Free quota: 1 million tokens per model, valid 90 days from account activation, model release or approval — Singapore only, not available in other regions.
Qwen-Flash (cheapest current tier): the exact current Qwen3.8-Flash price is not clear from the cited primary pricing material. Multiple independent trackers give a reasonably convergent range of $0.10–$0.15/M input, $0.40–$0.47/M output. It is excluded from the headline comparison because the primary rate is not established; it would be the cheapest Qwen tier and materially change any "Qwen vs DeepSeek Flash" framing.
2. Normalized Workload Cost (CALCULATION)
At the brief's three volumes, using each lab's cheapest confirmed non-legacy flagship-adjacent tier (Kimi K2.7 Code shown separately since it is a different model class than K3):
| Workload | Kimi K3 | Kimi K2.7 Code | DeepSeek Flash (off-peak) | DeepSeek Flash (peak) | DeepSeek Pro (off-peak) | Qwen Plus (Singapore, captured display rate) | Qwen Max (Singapore) |
|---|---|---|---|---|---|---|---|
| 10M in + 2M out | $60.00 | $17.50 | $2.70 | $5.40 | $10.56 | $7.20 | $32.00 |
| 100M in + 20M out | $600.00 | $175.00 | $27.00 | $54.00 | $105.60 | $72.00 | $320.00 |
| 1B in + 200M out | $6,000.00 | $1,750.00 | $270.00 | $540.00 | $1,056.00 | $720.00 | $3,200.00 |
None of these figures include cache hits. Real-world cost depends heavily on cache-hit rate, and the three labs' cache economics differ sharply: DeepSeek's cache-hit input is 2–4% of cache-miss price (the deepest discount of the three); Kimi's is 10%; Qwen's is 10–20% depending on model, with several current-generation exceptions the console must be checked for individually. A workload with a 90%+ cache-hit rate (a coding agent resending repository context, for instance) can bring effective input cost on any of the three down by an order of magnitude — third-party OpenRouter data cited elsewhere suggests real-world K3 traffic runs around 92% cache hits, which would cut K3's effective input cost roughly in half of what the headline $3.00/M implies, though this is a market-wide observation, not a Moonshot-published figure.
3. Hidden Costs and Structural Risk
- K3's always-on reasoning is the single biggest hidden-cost risk in this comparison. Every output token — including the model's internal reasoning trace, which must be sent back on subsequent turns for correct multi-turn behavior — bills at the full $15/M output rate. A harness that strips
reasoning_contentbetween turns will silently degrade the model's behavior; one that doesn't will pay for the full trace every turn. - DeepSeek's peak/off-peak clock is a scheduling optimization, not a discount you opt into. A workload that happens to run during the two daily UTC peak windows pays exactly double what the same workload run at 3am UTC on a Tuesday would. Schedulable, non-interactive workloads (batch summarization, offline evaluation) can be deliberately timed to avoid peak hours; interactive user-facing traffic cannot.
- Qwen's region-dependent pricing is the most consequential hidden cost of the three, because it is invisible unless you specifically check: the same model ID, called from a different Alibaba Cloud region setting, can cost 2–6× more or less. A team that builds a cost model against Singapore rates and later deploys to a US or EU endpoint for latency or compliance reasons will see a materially different bill without any model or plan change.
- Qwen's input-length tiering means a single long request can jump the entire request's price, not just the tokens above the threshold — per Alibaba's own documentation, a 100K-token request on a 32K/128K-tiered model is billed entirely at the 32K–128K tier rate, not a blended rate. This rewards keeping requests under a tier boundary far more than DeepSeek's or Kimi's flat-rate structures do.
- K3's open weights are not a free self-hosting escape hatch for most teams. The $20M-revenue MaaS threshold matters only to resellers, but the hardware floor (multiple high-end GPUs, even at reduced precision) puts genuine self-hosting out of reach for most individual developers and small teams regardless of licensing.
- None of the three publishes a stable "current" price with confidence beyond the immediate term. DeepSeek's own pricing page explicitly warns prices may change and to check regularly; Qwen's Plus/Max tiers are explicitly marked as carrying "limited-time" promotional discounts on top of list prices; Kimi's K2 lineup was repriced and partially retired (K2.5) within the last two months of this research window.
4. Commercial and Regional Access
All three offer standard commercial API terms for hosted use (no separate commercial license required to call the API, as distinct from self-hosting open weights — a distinction covered in NotCheapAI's FLUX self-hosting article for a different provider). Regional availability and compliance posture differ meaningfully:
- DeepSeek: China-hosted infrastructure per its own documentation; no explicit region-selection mechanism was found on its pricing page (single global rate, adjusted only by time of day).
- Qwen: explicit multi-region deployment (Singapore, China, Hong Kong, EU/Frankfurt, US/Virginia, Japan) with materially different prices and, in most non-Singapore regions, time-of-day discounts layered on top — the most flexible regional posture of the three, but also the hardest to budget against without picking a region first.
- Kimi/Moonshot: data residency reports conflict (Singapore vs. China), and the cited public materials do not settle the issue. Treat it as an open compliance question for regulated workloads.
5. BUILD & EARN Scenario: A Document-Summarization API Wrapper
MODELING ASSUMPTION: a service charges $0.02 per summarized document (avg. 3,000 input tokens, 300 output tokens per doc), processing 500,000 documents/month = $10,000 revenue. That workload is 1.5B input tokens + 150M output tokens/month. Metric: model-cost gross margin (excludes hosting, retries, review — not profit).
| Route | Token cost (CALCULATION) | Share of $10,000 | Model-cost gross margin |
|---|---|---|---|
| DeepSeek Flash, off-peak | 1,500×$0.15 + 150×$0.60 = $315.00 | 3.2% | 96.8% |
| Qwen Plus, Singapore captured display rate | 1,500×$0.40 + 150×$1.60 = $840.00 | 8.4% | 91.6% |
| Kimi K2.7 Code | 1,500×$0.95 + 150×$4.00 = $2,025.00 | 20.3% | 79.7% |
| Kimi K3 | 1,500×$3.00 + 150×$15.00 = $6,750.00 | 67.5% | 32.5% |
At this volume, the choice of flagship model is the single largest lever on margin — K3 alone would consume two-thirds of revenue on token cost before any caching benefit, while DeepSeek's off-peak Flash rate leaves the business with a 96.8% gross margin on the same workload. Whether K3's quality improvement over cheaper tiers justifies that gap depends entirely on the task, which this article does not assess.
FAQ
Is Kimi K3 still cheap? Not relative to DeepSeek or Qwen's budget tiers — at $3/$15 per million tokens it's priced in the same range as Western mid-tier models, a genuine break from the K2 generation's positioning.
Which is cheapest for high-volume production? DeepSeek's Flash model at off-peak hours, by a wide margin over every other tier examined here — but only for workloads that can tolerate China-region hosting and the time-of-day billing structure.
Does Qwen have one price? No — the same model ID can cost 2–6× more or less depending on which Alibaba Cloud region you call it from, and most models add a further price step once a single request exceeds 32K–256K input tokens.
Can I self-host any of these for free? Kimi K2.6/K2.7 Code and K3 (post-July 27, 2026) ship open weights; most Qwen models are open-weight under Apache 2.0. DeepSeek's current V4.1-Flash/V4-Pro generation's weight-release status is unclear from the sources cited here. Self-hosting any of them at production scale requires real GPU infrastructure regardless of license terms.