NotCheapMAKE EVERY CREDIT COUNT
10providers compared
$70-$2,700monthly workload C range
80%cache-hit scenario modeled

Chinese AI Price War 2026: Real API Cost Across Ten Providers

Verification date: September 20, 2026. Token prices, promotions and peak/off-peak schedules change often.


DIRECT ANSWER

The short answer

The Chinese AI price war is no longer about one headline rate. Across ten providers, three things now decide what you pay: how cheaply cached input is billed, what time of day (or context length) you call the model, and which region's price list you are reading.

  • The spread between large-model prices is huge. For a 400M-input + 100M-output month with no caching, the cheapest of the six pro-tier models in our tables (Kimi K3, Qwen3.8-Max, GLM-5.3, DeepSeek V4-Pro, Tencent Hy4 preview, MiMo-V2.5-Pro) costs $261 (MiMo-V2.5-Pro) and the most expensive $2,700 (Kimi K3). That is about a 10× gap in list price ($2,700 ÷ $261 ≈ 10.34×). It says nothing about which model does your job better, and we make no quality comparison.
  • Caching reshuffles the ranking. Cached input costs between 0.8% and 20% of the normal input price depending on the provider. With 80% cache hits, Xiaomi's flagship falls to $26 for a 100M-in/20M-out job, below Step 3.7 Flash ($30.20) and DeepSeek's peak-priced Flash ($30.48).
  • Some prices depend on the clock. DeepSeek charges twice as much at weekday peak hours as off-peak. MiniMax doubles its rate above 512K input tokens.
  • Two cells could not be confirmed. Alibaba's cheap qwen3.8-flash and ByteDance's international BytePlus text models have no confirmed numeric rate, so they are not in any quantitative comparison here. We never substitute mainland Chinese prices for international ones.
  • Cheap per token is not cheap per task. Reasoning models can spend many more output tokens, and output is where the money goes.

1. The ten providers and their current models

ProviderModel used hereInput / cache hit / output ($ per 1M)Context
MoonshotKimi K3 (international)$3.00 / $0.30 / $15.001,048,576
AlibabaQwen3.8-Max$2.00 / $0.25 / $6.001M
Zhipu / Z.aiGLM-5.3$1.40 / $0.26 / $4.40about 1M
DeepSeekdeepseek-v4-pro (V4-Pro-0813)Peak $1.32 / $0.044 / $3.96; off-peak $0.66 / $0.022 / $1.981M
TencentHy4 preview$0.834 / $0.042 / $2.5011M
XiaomiMiMo-V2.5-Pro$0.435 / $0.0036 / $0.871M
MiniMaxMiniMax-M3 (up to 512K input)$0.30 / $0.06 / $1.20Two price tiers (≤512K and >512K input)
StepFunStep 3.7 Flash$0.20 / $0.04 / $1.15256K
BaiduERNIE 5.1Mainland only: ¥4 in / ¥18 out (up to 32K input); ¥6 / ¥22 (32K-128K)128K; excluded from USD tables
ByteDanceDoubao / Seed text model via BytePlus (international)Not availableNot confirmed

Budget-tier models from the same providers: DeepSeek deepseek-flash (V4.1-Flash): $0.15 / $0.003 / $0.60 off-peak, $0.30 / $0.006 / $1.20 peak. GLM-5.3-Flash: $0.15 / $0.03 / $0.50. MiMo-V2.5: $0.14 / $0.0028 / $0.28. Step 3.5 Flash: $0.10 / $0.02 / $0.30.

Pricing caveats:

  • Kimi K3. The Kimi Open Platform's mainland list is ¥20 cache-miss input, ¥2 cache-hit input and ¥100 output per 1M tokens, with a 1,048,576-token context. The international dollar rate above is a separate listing. At an assumed ¥7.2 per US$1 it is about 8% above the mainland equivalent; keep the two price lists separate.
  • Qwen. qwen3.8-max is listed at $2.00 input and $6.00 output. Published reports agree, and one reports a flat rate across the full 1M context. Alibaba's cache documentation says the cache-hit price is not a fixed percentage of input and points to the console, so the $0.25 cache figure should be treated as a reported rate rather than a published fixed formula.
  • Tencent. Hy4 preview's identity and specs (770B total, 49B active, open weights) are documented. Public pricing evidence is less direct than for the other models in this table.
  • Baidu. ERNIE 5.1's dollar prices are not confirmed. The mainland yuan list applies to mainland accounts, and one review notes Baidu's international catalog lists ERNIE 5.0, not 5.1. We do not convert it into a USD price.
  • MiniMax. The page shows M3 with a "permanent 50% off" tag. The $0.30 / $1.20 above is the discounted rate; the struck-through list rate is $0.60 / $2.40.

The two unresolved cells:

  • qwen3.8-flash: Alibaba's model page confirms it is current, has a 1M-token context and supports context caching. Its exact input, output and cache rates are not published in the pricing material cited here. Third-party trackers disagree, so it is excluded from the comparison.
  • BytePlus international Doubao/Seed: Official documentation confirms the international service, its Singapore and EU regions, and the current catalogue, but the numeric international pay-as-you-go table is not available in the cited material. Volcano Engine's mainland rates are a different price list and are not used here.

2. Real workloads

Formula (prices per 1M tokens, volumes in millions):

Cost = (Input × HitShare × P_hit) + (Input × (1 − HitShare) × P_miss) + (Output × P_out)

Workload D assumes 80% of input hits the cache (a MODELING ASSUMPTION; real hit rates depend on how stable your prompts are). Peak-priced rows use DeepSeek's peak rates.

ModelA: 1M in + 250K outB: 10M in + 3M outC: 400M in + 100M out (monthly)D: 100M in, 80% cached + 20M out
Step 3.5 Flash$0.175$1.90$70.00$9.60
MiMo-V2.5$0.21$2.24$84.00$8.62
GLM-5.3-Flash$0.275$3.00$110.00$15.40
DeepSeek V4.1-Flash, off-peak$0.30$3.30$120.00$15.24
Step 3.7 Flash$0.4875$5.45$195.00$30.20
MiniMax-M3 (≤512K)$0.60$6.60$240.00$34.80
DeepSeek V4.1-Flash, peak$0.60$6.60$240.00$30.48
MiMo-V2.5-Pro$0.6525$6.96$261.00$26.39
DeepSeek V4-Pro, off-peak$1.155$12.54$462.00$54.56
Tencent Hy4 preview$1.459$15.84$583.70$70.06
DeepSeek V4-Pro, peak$2.31$25.08$924.00$109.12
GLM-5.3$2.50$27.20$1,000.00$136.80
Qwen3.8-Max$3.50$38.00$1,400.00$180.00
Kimi K3 (international)$6.75$75.00$2,700.00$384.00

CALCULATION on the rates in Section 1. Worked example, workload C on Kimi K3: 400 × $3.00 + 100 × $15.00 = $1,200 + $1,500 = $2,700. Workload D on MiMo-V2.5-Pro: 80 × $0.0036 + 20 × $0.435 + 20 × $0.87 = $0.288 + $8.70 + $17.40 = $26.39. Qwen3.8-Max's workload D uses the secondary-reported $0.25 cache price; Tencent Hy4's rows rest on secondary-classified pricing; the Kimi K3 rows use the international rate discussed above.

How to read it. The price tiers are not quality tiers. Rows are grouped only by price. We also did not adjust for how many tokens each model needs to finish a task (Section 5).


3. The caching war

Every provider here discounts cached input, but by very different amounts (cache-hit price as a share of the normal input price, CALCULATION):

ModelCache hit ÷ cache miss
MiMo-V2.5-Pro0.8%
DeepSeek V4.1-Flash2.0%
MiMo-V2.52.0%
DeepSeek V4-Pro3.3%
Tencent Hy4 preview5.0%
Kimi K310%
Qwen3.8-Max12.5% (secondary cache price)
GLM-5.318.6%
MiniMax-M3, StepFun, GLM-5.3-Flash20%

Three consequences:

  1. A high hit rate rewards a small ratio. In workload D, caching cuts Xiaomi's MiMo-V2.5-Pro bill by 57% ($60.90 to $26.39) and DeepSeek V4-Pro's by 48%. It cuts Step 3.7 Flash's by 30%, because its 20% ratio saves less and most of its bill is output.
  2. Rankings flip. On workload D, MiMo-V2.5-Pro ($26.39) beats Step 3.7 Flash ($30.20), DeepSeek's peak-priced Flash ($30.48) and MiniMax-M3 ($34.80), though it costs more on uncached input than three of them.
  3. Output still dominates. In workload D on MiMo-V2.5-Pro, $17.40 of the $26.39 is output, and caching cannot discount it. Kimi K3's output rate ($15) is 17× Xiaomi's flagship rate ($0.87).

Cache hits are not guaranteed anywhere. StepFun, for instance, says it cannot force every request to hit, and Alibaba's cache price is not published as a simple percentage. Measure your own hit rate before budgeting on a discount.


4. The clock, the context and the region

DeepSeek's peak and off-peak. Its Flash and Pro rates are doubled at peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC (09:00–12:00 and 14:00–18:00 Beijing time), Monday to Friday, excluding Chinese public holidays; everything else is off-peak. That is 7 peak hours on each weekday, or 35 of the 168 hours in a week (about 20.8%). Off-peak is therefore 133 of 168 weekly hours (about 79.2%): 17 hours on each weekday and all 24 hours on weekends and Chinese public holidays, so weeks with a public holiday have fewer peak hours. If your traffic were spread evenly across the week (a MODELING ASSUMPTION), your blended price would be about 1.21× the off-peak rate. Workload C on V4.1-Flash, uncached: $120 off-peak, $240 all-peak, about $145 evenly spread. On V4-Pro: $462, $924 and about $558. Scheduling batch work off-peak halves the bill.

MiniMax's long-context tier. Above 512K input tokens, MiniMax-M3 charges $0.60 input, $0.12 cache read and $2.40 output, twice the shorter tier. Workload C would cost $480, not $240. Independent reporting describes Alibaba's Qwen3.8-Max rate as flat across its 1M context; check the Alibaba console for the current account-specific price.

Region matters. Prices differ by route, and we keep them separate:

  • Kimi's mainland list of ¥20 / ¥2 / ¥100 is not a global USD price. At an illustrative ¥7.2 per US$1 (MODELING ASSUMPTION), it is about $2.78 / $0.28 / $13.89, versus $3.00 / $0.30 / $15.00 internationally.
  • Alibaba's international (Singapore) endpoint is priced differently from Beijing, which one secondary source reports as 60–70% cheaper for some models. Mainland endpoints can carry different account, data-residency and access conditions.
  • Baidu's ERNIE 5.1 is listed only in yuan for mainland accounts. We do not turn it into an official dollar price.

5. Cheap per token is not cheap per task

Token prices compare inputs and outputs. They do not compare how many tokens each model needs to do the same work. We make no claim about which model is better, but four things can move your real bill:

  • Reasoning tokens. Reasoning models produce thinking tokens before the answer. Z.ai states no separate reasoning surcharge (a secondary report of its pricing page), which means thinking is billed at the output rate. Check each provider's rule.
  • Effort settings. GLM-5.3 is reported as reasoning-always-on with a default of maximum effort. StepFun's Step 3.7 Flash offers low, medium and high; DeepSeek's V4 models support thinking and non-thinking modes.
  • Vendor token-efficiency claims. Xiaomi says its flagship uses roughly 40–60% fewer tokens per agent task than some peers, and StepFun's own model card says Step 3.5 Flash currently needs longer trajectories than Gemini 3.0 Pro. Both are vendor statements. Test on your own tasks.
  • Host prices are not vendor prices. Third-party hosts can quote rates far from the vendor's list. OpenRouter's headline price for GLM 5.2 (about $0.49 input) was roughly a third of Z.ai's own $1.40, reflecting a cheaper third-party host, and host serving configurations can differ. We use vendor lists except where labeled.

6. Build & Earn: a cost-routed AI support-agent product

Scenario (all MODELING ASSUMPTIONs): a SaaS product charges $99 per customer per month. Each customer's usage is 20M input tokens (70% cache hits) and 3M output tokens a month. The model-cost formula is the same as above.

ModelModel cost per customerModel-cost share of $99Model-cost gross margin
Kimi K3$67.2067.9%32.1%
Qwen3.8-Max$33.5033.8%66.2%
GLM-5.3$25.2425.5%74.5%
DeepSeek V4-Pro, peak$20.4220.6%79.4%
Tencent Hy4 preview$13.1013.2%86.8%
DeepSeek V4-Pro, off-peak$10.2110.3%89.7%
MiniMax-M3$6.246.3%93.7%
MiMo-V2.5-Pro$5.275.3%94.7%
DeepSeek V4.1-Flash, off-peak$2.742.8%97.2%
Step 3.5 Flash$1.781.8%98.2%

CALCULATION. Example, Kimi K3: 20 × 0.7 × $0.30 + 20 × 0.3 × $3.00 + 3 × $15.00 = $4.20 + $18.00 + $45.00 = $67.20.

Routing. A product that sends 10% of traffic to Kimi K3 and 90% to MiMo-V2.5-Pro would pay about $11.46 per customer (11.6% of price), against $67.20 for Kimi alone. Sending 20% to GLM-5.3 and 80% to DeepSeek V4.1-Flash off-peak would cost about $7.24 (7.3%). Whether the cheaper model can handle 80–90% of your traffic is a quality question this article does not answer.

What this does and does not show. "Model-cost gross margin" here counts only model API fees. It is not business profit. It excludes hosting, support, sales, payments, engineering and the cost of retries. At 100 customers, the model bill ranges from about $274 a month (DeepSeek Flash off-peak) to $6,720 (Kimi K3), so model choice and cache hit rate are among your largest levers, but not the only ones.