NotCheapMAKE EVERY CREDIT COUNT

Kimi K3 Isn't the Cheap Chinese Model You Expected — Here's What It Actually Costs

Executive summary

Kimi K3, Moonshot AI's flagship model, is not priced like DeepSeek or Qwen's budget tiers. Its official API rate — confirmed directly on Moonshot's own pricing documentation — is $3.00 per million input tokens and $15.00 per million output tokens, with a $0.30 per million cache-hit rate. That's the same pricing tier as Claude Sonnet, not the sub-dollar rates that "cheap Chinese model" coverage might lead you to expect. What actually makes K3 economically interesting is its 1,048,576-token context window combined with that cache-hit discount: workloads that repeatedly reuse a large context can push effective cost down toward $0.30/M, while workloads that don't — short chats, one-off prompts, or reasoning-heavy tasks with long outputs — pay close to full price on a model whose output rate is 5x its cache-hit input rate. This article works through the actual math rather than the sticker price.

What is Kimi K3

Kimi K3 is Moonshot AI's flagship large language model, launched as a hosted API on July 16, 2026, with full weights published on Hugging Face and GitHub on July 27, 2026. Per Moonshot's own technical materials (reported in detail by Unite.AI's coverage of the license release):

  • 2.8 trillion total parameters, with 104 billion active per token — a mixture-of-experts (MoE) design Moonshot calls Stable LatentMoE, routing each token to 16 of 896 experts. Moonshot describes it as the first open-weight model to reach the 3-trillion-parameter class.
  • 1,048,576-token context window — confirmed directly on Moonshot's own pricing page.
  • Native multimodal (vision) input through a vision encoder reported as "MoonViT-V2," taking text and images through the same pipeline rather than bolting vision on separately. (This specific encoder name comes from a single detailed secondary source, not Moonshot's own model card directly — worth treating as reported rather than fully primary-confirmed.)
  • Architecturally, Kimi Delta Attention handles most of the model's layers, paired with a mechanism Moonshot calls Attention Residuals, which the lab reports yields roughly a 2.5x improvement in scaling efficiency over Kimi K2 — a claim independently repeated in AWS's own Bedrock launch announcement, which is a reasonable corroboration since AWS has no obvious reason to inflate a partner's efficiency claim.

On independent benchmarking, Artificial Analysis scored K3 at 57 on its Intelligence Index — third overall, roughly level with Claude Opus 4.8 and GPT-5.5, behind Claude Fable 5 and GPT-5.6 Sol. Its agentic benchmark numbers were reported as the stronger half of that result (1668 Elo on GDPval-AA v2, first place on AutomationBench-AA at 53%), with a cost-per-task figure of roughly $0.94 — about half of Opus 4.8's — a genuinely useful data point since it's cost-per-completed-task rather than cost-per-token, and it's Artificial Analysis's own measurement rather than Moonshot's.

API pricing — the real numbers

Moonshot's model pricing documentation (platform.moonshot.ai/docs/pricing, mirrored at platform.kimi.ai) lists:

ModelCache-hit inputCache-miss inputOutputContext window
kimi-k3$0.30 / M$3.00 / M$15.00 / M1,048,576 tokens
kimi-k2.7-code$0.19 / M$0.95 / M$4.00 / M262,144 tokens
kimi-k2.7-code-highspeed$0.38 / M$1.90 / M$8.00 / M262,144 tokens
kimi-k2.6$0.16 / M$0.95 / M$4.00 / M262,144 tokens

Kimi K2.5 does not appear in Moonshot's current official pricing table at all. A separate source (Morph's Kimi API guide) reports it and the older moonshot-v1 series were retired on August 31, 2026 and now return a 404 if called — consistent with its absence from the current live pricing page. Treat any "K2.5 pricing" you see elsewhere as describing a model that's no longer active; we found at least one aggregator still listing K2.5 at a price, which is stale.

There's no off-peak discount and no batch-pricing tier listed for K3 specifically (Moonshot's Batch API, where it exists, applies to K2.6 and K2.7 Code at 60% of standard rates — K3 is explicitly excluded from batch pricing per Moonshot's own documentation structure). Unlike DeepSeek, which runs a peak/off-peak clock, K3's rate is flat around the clock.

The built-in web search tool's per-call fee is listed in Moonshot's pricing documentation. K3 supports a native $web_search tool, and search results returned into context bill as ordinary input tokens on the next call. Moonshot's tool-pricing page (platform.kimi.ai/docs/pricing/tools) states: "When you add the $web_search tool in tools and receive a response with finish_reason = tool_calls and tool_call.function.name = $web_search, we charge a fee of $0.005 for the $web_search call." This differs from the $0.015 figure in a third-party reseller's (EmpirioLabs) documentation; $0.005 is the listed Moonshot rate for this specific tool. The same page documents standalone REST endpoints - /v1/tools/search (Web Search Basic, $0.002/call) and /v1/tools/search_pro (Web Search Pro, $0.003/call) - which bill differently from the in-chat tool.

The comparison that actually matters: K3's $3/$15 is meaningfully below Claude Opus 5's $5/$25 — about 40% cheaper on both sides — but it's in the same tier as Claude Sonnet, not competing with DeepSeek's or Qwen's cheapest models, which run well under $1/M on both input and output. If you came to K3 expecting DeepSeek-level pricing, you're comparing it to the wrong class of model.

Consumer subscription plans (kimi.com)

Separate from the metered API, Moonshot sells a consumer subscription for the Kimi web/mobile app, Kimi Code CLI, and agent features. Independent trackers report a free Adagio tier, then Moderato ($19/mo), Allegretto ($39/mo), Allegro ($99/mo), and Vivace ($199/mo), each with a modestly discounted annual option. Confirm current consumer-plan terms before subscribing. Reports indicate all tiers get K3 access in the app; the agent-credit pool, concurrency, and availability of the 1M-token context in Kimi Code vary by tier, with multiple sources placing the latter at Allegro and above. This subscription does not include API credit - it's a separate product from everything priced in the table above.

Two caveats matter. Some reports say these tiers briefly showed as "sold out" in late July 2026 before becoming consistently available again from August onward. A separate report describes a China-region or Kimi-Code-specific variant priced in CNY under different tier names and counts (Andante/Moderato/Allegretto/Allegro at ¥49/¥99/¥199/¥699); public information does not establish how that ladder relates to the USD tiers. A separate claim of a $600/year/seat "Kimi Business" tier appears in a single report without independent corroboration, so treat it as unconfirmed.

Context caching: why the cache-hit rate is the number that matters

Moonshot bills cache-hit input at $0.30/M against a $3.00/M cache-miss rate — a 10x discount, applied automatically (Moonshot's documentation and multiple independent write-ups describe context caching as standard/automatic on the official API, not an opt-in feature you have to configure).

Why this changes the economics: the first time you send a large block of context — a codebase, a long document, a big system prompt — you pay the full $3.00/M cache-miss rate to ingest it. Every subsequent call that reuses the same prefix pays $0.30/M for the reused portion instead. For a workload that repeatedly analyzes the same large document or codebase across many calls, effective input cost trends toward $0.30/M rather than $3.00/M. One source (Codersera) reports Moonshot claims its official API exceeds a 90% cache-hit rate in coding workloads specifically — that's a vendor claim tied to Moonshot's own serving infrastructure (a stack it calls Mooncake), not something we independently verified, so treat it as reported rather than confirmed.

The caveat that matters more than the discount: caching only helps when there's a stable, repeated prefix to cache. A workload made of genuinely distinct, one-off prompts — no shared context between calls — never benefits from the $0.30 rate at all. Every token in that kind of workload is a cache miss, billed at full $3.00/M input. This is the single most important qualifier on any "K3 is cheap because of caching" claim: the caching discount is real, but it's conditional on your actual usage pattern, not a blanket discount on the model.

Reasoning cost: the part that can quietly dominate your bill

Kimi K3 runs with always-on reasoning. Per one detailed technical guide to the Kimi API (Morph), the model exposes a reasoning_effort parameter with levels low/high/max, defaults to max, and there is no way to fully disable thinking. Separately, Moonshot's own release materials describe K3 as trained in what the lab calls "preserved thinking history mode" — meaning a harness needs to pass the model's complete prior reasoning content and tool calls back into the conversation on every turn, not just the visible reply, or output quality degrades.

The economic consequence is direct: reasoning tokens are billed as output tokens, at K3's $15.00/M output rate — 5x the cache-hit input rate and 5x the cache-miss input rate's ratio doesn't even apply here, since output has no cached discount at all. For any task where the model reasons at length before answering, the reasoning tokens themselves — not the visible answer — can be the majority of what you're actually paying for. This is precisely why the executive summary above singles out "reasoning-heavy tasks with long outputs" as a weak economic fit: there's no cache discount available to soften that side of the bill, and it's a fixed behavior of the model, not a setting you can turn off to save money.

Kimi K3 on Amazon Bedrock

AWS's "What's New" announcement and Machine Learning Blog post state that Kimi K3 became generally available on Amazon Bedrock on September 18, 2026:

  • Kimi K3 is the first open-weight model on Amazon Bedrock to support explicit prompt caching.
  • Available in all AWS Regions where Bedrock operates, via cross-Region inference.
  • AWS's Bedrock model card for Kimi K3 (docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html) lists both cross-Region inference profiles: "Kimi K3 is available through US Geo cross-Region inference (using the us.moonshotai.kimi-k3 profile, which routes requests only among US-geography Regions to respect US data residency) and Global cross-Region inference (using the global.moonshotai.kimi-k3 profile, which routes to any supported commercial AWS Region worldwide)."
  • On data handling: AWS's own Bedrock materials state, generally for third-party models, that data stays inside the AWS boundary, isn't shared with the model provider, and isn't used for training — which is the substance of "zero data retention" even though we did not find AWS's own Kimi K3 launch pages using that exact phrase. Separately, Moonshot's own platform documentation describes Zero Data Retention as an enterprise-tier feature of its own Kimi Open Platform generally (deleting prompts and responses immediately after a response is returned) — a real Moonshot feature, but described independently of the Bedrock launch specifically, so we're not merging the two into one claim.
  • Runs within the same AWS security boundary as proprietary Bedrock models — data stays inside the AWS data boundary, isn't shared with Moonshot, and isn't used for training.

On the "global is cheaper" claim: this is real, but it's a general Bedrock platform mechanic, not something specific to Kimi K3's launch. AWS's own Bedrock documentation states plainly: "Global cross-Region inference offers approximately 10% savings on both input and output token pricing compared to geographic cross-Region inference." Bedrock maintains a pricing catalog separate from Moonshot's direct-API rate card; check AWS's current pricing page for the exact Kimi K3 rate before budgeting, and don't assume it's identical to Moonshot's direct-API $3/$15.

Open-weight licensing: what it actually says

Kimi K3's weights are released under a bespoke "Kimi K3 License," not a standard OSI-approved open-source license and not plain MIT, despite reading like MIT for most of its length. Per a detailed breakdown of the license terms (Unite.AI, which reviewed the license text directly):

  • The license grants, free of charge, the right to use, copy, modify, distribute, sublicense, sell, deploy, fine-tune, and build derivative works from the weights.
  • Two conditions attach above a revenue line — and neither is a revenue-share arrangement. A "Model as a Service" business (defined as giving third parties inference or fine-tuning access with meaningful control over inputs, parameters, or training data) must sign a separate agreement with Moonshot once its revenue, combined with its affiliates, passes $20 million over any consecutive 12-month period. What that separate agreement actually requires — including whether it involves anything resembling a revenue share — is not publicly disclosed in the license text itself, so no percentage can be stated. No primary-source support is available for a "30% revenue share" figure attached to Kimi K3, so treat that claim as unverified.
  • Separately, any commercial product or service with more than 100 million monthly active users, or more than $20 million in monthly revenue, must display "Kimi K3" prominently in its interface — a branding requirement, not a payment.
  • Neither condition applies to purely internal use, or to access through Moonshot's own products and certified inference partners.

The practical upshot: self-hosting is legally straightforward for the overwhelming majority of use cases — internal deployments and small-to-mid-scale products owe Moonshot nothing. It's a narrow, high-revenue slice of commercial MaaS resellers that needs to talk to Moonshot directly.

Financially, "open-weight" is not "free to run." The Hugging Face checkpoint spans roughly 1.56 TB across 96 safetensors shards, and Moonshot's own deployment guidance recommends supernode configurations of 64 or more accelerators, using vLLM, SGLang, or TokenSpeed as inference engines. This is not a model you casually self-host on a single GPU box to save API costs — the hardware floor alone puts self-hosting out of reach for anyone below serious enterprise scale.

Real cost examples

All four figures below are independently recalculated against Moonshot's confirmed rate table above, not assumed from the brief that produced this article's outline:

ScenarioTokensCalculationCost
Short chat (K3)2,000 input + 500 output2,000×$3/M + 500×$15/M$0.0135
Coding task, 70% cached (K2.7 Code)50,000 input (35,000 cached / 15,000 uncached) + 2,000 output(35,000×$0.19/M + 15,000×$0.95/M) + 2,000×$4/M$0.0289
Long document, fully cached (K3)950,000 cached input + 5,000 output950,000×$0.30/M + 5,000×$15/M$0.36
Monthly workload, 50% cached (K3)400M input (200M cached / 200M uncached) + 100M output(200M×$0.30/M + 200M×$3/M) + 100M×$15/M$2,160

The gap between the coding-task example and the long-document example is the whole article in miniature: both use a large context, but the coding task's 30% cache-miss share and small output keep it under three cents, while the long-document example — despite being almost entirely cached — still costs 25x more, purely because the 950,000-token document dwarfs the 50,000-token coding context in absolute size. Caching reduces the rate, not the fact that a very large context is still a lot of billed tokens even at a discount.

Practical recommendation

Good fit for Kimi K3:

  • Repeated analysis of the same long document, contract, or codebase across many calls, where the cache-hit discount compounds over volume.
  • Legal, financial, or research workflows built around a large, stable reference context that changes infrequently.
  • Enterprise use cases already on AWS, where Bedrock's explicit prompt caching and existing security/governance boundary matter more than shaving pennies off the per-token rate.

Weak fit for Kimi K3:

  • Short, casual chat with no repeated context — every token is a cache miss at full price, and there's a cheaper model in almost every other provider's lineup for this.
  • One-off prompts with no reuse pattern at all — the caching mechanic that's supposed to make K3 economical never engages.
  • Output-heavy or reasoning-heavy tasks — the always-on thinking mode bills reasoning as output tokens at $15/M with no cache discount available on that side, and this is a structural feature of the model, not a setting you can adjust.

Should you use Kimi K3?

If your workload genuinely reuses a large context — a codebase, a growing document, an enterprise knowledge base — Kimi K3's 1M-token window and $0.30/M cache-hit rate can make it a legitimately strong option, and one that undercuts Claude Opus 5 by a real margin. If your workload is short chats, one-off prompts, or long reasoning chains that produce a lot of output, K3's Sonnet-class sticker price applies in full, the caching discount you were promised never shows up, and a cheaper model — or a different one entirely — is very likely the better economic choice. The mistake this article is written to prevent is assuming "open-weight Chinese model" means "cheap by default." On the numbers, it doesn't. It means "cheap under a specific, checkable usage pattern" — and it's worth checking which pattern you actually have before committing production spend.

FAQ

Is Kimi K3 cheaper than Claude? Cheaper than Claude Opus 5 (about 40% less on both input and output) — but priced in the same tier as Claude Sonnet, not against Claude's or anyone else's budget models.

Does Kimi K3 have a free tier? The API itself is pay-per-token, with no free tier reported. Moonshot's consumer-facing kimi.com app is reported as free to use by multiple sources, which is a separate product from the metered API this article covers.

Is Kimi K3 open source? It's open-weight under a bespoke "Kimi K3 License" that permits free commercial use for the large majority of use cases, with a revenue-gated separate-agreement requirement for large-scale Model-as-a-Service resellers above $20M/12 months — not a standard OSI open-source license, and not the "30% revenue share" some secondary coverage has claimed, which the cited public sources do not confirm.

Can I self-host Kimi K3 to avoid API costs? Legally, for most use cases, yes. Financially, the model's recommended deployment footprint (64+ accelerators, a 1.56 TB checkpoint) puts real self-hosting economics out of reach below serious enterprise infrastructure scale — this isn't a cost-saving move for a small team.

Why does caching matter so much for Kimi K3 specifically? Because its 1M-token context window makes it attractive for exactly the kind of large-context, repeated-reuse workloads where a 10x cache discount compounds the most — but that discount is entirely conditional on your workload actually having a stable, reused prefix.

For the wider market context, see our Chinese AI API price comparison.