NotCheapMAKE EVERY CREDIT COUNT

Xiaomi MiMo API vs Self-Hosting: Does MIT-Licensed AI Actually Save Money?

Verification date: September 19, 2026. Prices and promotions move.


DIRECT ANSWER

The short answer

Xiaomi published MiMo-V2.5-Pro's weights under the MIT license. Anyone can download a 1.02-trillion-parameter model and run it. That does not make running it cheap.

  • The hosted price is very low. Xiaomi's own API charges $0.435 per million input tokens, $0.87 per million output tokens and $0.0036 per million cached-input tokens. A month of 400M input + 100M output tokens costs $261.
  • The smallest documented self-hosting setup is one eight-GPU node (H200 or B200 class). Rented around the clock at July 2026 cloud rates, that is roughly $10,000–$31,000 a month before any engineering time.
  • Self-hosting wins on price only in a narrow corner: input-heavy work, little prompt caching, cheap GPU capacity and enough sustained utilization to clear break-even. In our models the break-even utilization ranges from about 26% to well over 100%. Over 100% means self-hosting never wins on price.
  • Non-price reasons are real. If the data cannot leave your infrastructure, or you need to fine-tune the weights, the API's price is beside the point.

The rest of this article shows the arithmetic and separates published rates from modeled assumptions.


1. What MiMo-V2.5-Pro actually is

ItemDetails
Model nameMiMo-V2.5-Pro (Xiaomi API ID mimo-v2.5-pro; Hugging Face XiaomiMiMo/MiMo-V2.5-Pro; OpenRouter xiaomi/mimo-v2.5-pro)
ArchitectureMixture-of-Experts, 1.02T total / 42B active parameters; 384 routed experts, 8 used per token; 70 layers
ContextUp to 1M tokens (the Base variant is 256K)
WeightsFP8 (E4M3) mixed precision; the Hugging Face repo totals about 1.03 TB
ModalityText
DatesXiaomi's open-source announcement is dated April 27, 2026. OpenRouter lists an April 22, 2026 release for API access
StatusXiaomi calls it its most capable model to date. Xiaomi has since opened an invite-only beta of newer "MiMo-X" preview models; their relationship to V2.5-Pro is not documented

For scale, Xiaomi's own model-card table places it next to Kimi-K2 (1.04T total / 32B active) and DeepSeek-V4-Pro (1.6T total / 49B active). It is a trillion-class model.

The naming trap: V2.5 is not V2.5-Pro

Xiaomi ships two similarly named open-weight models. MiMo-V2.5 is a 310B-total / 15B-active omnimodal model. MiMo-V2.5-Pro is the 1.02T text flagship. Their list prices on Xiaomi's overseas rate card differ by exactly 3.1×:

ModelInput (miss)Output
mimo-v2.5$0.14$0.28
mimo-v2.5-pro$0.435$0.87

Before you compare any two AI prices, check that the model IDs match character for character.


2. The license: MIT, with conditions

The Hugging Face model card declares license: mit, and Xiaomi's official announcement on X says the MiMo-V2.5 series is MIT-licensed and open for commercial deployment, continued training and fine-tuning without extra authorization. Xiaomi's own blog calls it a "permissive license" without naming it.

Two precision notes for readers:

  1. Say it this way: commercial use is permitted under the MIT license terms. MIT still requires that the copyright notice and permission notice travel with copies or substantial portions of the software. It also disclaims warranty. "No restrictions" would be wrong.
  2. The MIT license covers the weights, not the hosted service. Using Xiaomi's API, or any third-party host, is governed by that provider's own terms.

The model repository's file listing has no standalone LICENSE file, so the license is declared through the model-card metadata and Xiaomi's statement. One third-party documentation page (the SGLang cookbook) lists the license as Apache 2.0. That conflicts with Xiaomi's own materials and may be a documentation error.


3. What hosted access costs today

3.1 Xiaomi's own rate card

Xiaomi publishes a pay-as-you-go rate card on its API platform docs (page stamped "Update Time August 06, 2026"). Overseas prices are in US dollars per million tokens.

mimo-v2.5-pro, overseasPrice per 1M tokens
Input, cache miss$0.435
Input, cache hit$0.0036
Output$0.87
Cache writeFree "for a limited time"
Web search plugin (overseas)$5 per 1,000 calls, billed separately

Artificial Analysis independently lists Xiaomi at $0.43 input / $0.87 output, and its published blended price of $0.18 per million (7:2:1 cache-hit/input/output mix) matches this rate card when you work it out (0.7 × $0.0036 + 0.2 × $0.435 + 0.1 × $0.87 ≈ $0.177)., CALCULATION.

The price has fallen fast. At launch, VentureBeat reported $1.00 input / $3.00 output for prompts up to 256K, with a higher tier up to 1M and a $0.20 cache-hit price. The current page shows a single rate set. That is a 56% cut on input, 71% on output and 98% on cache hits versus launch. VentureBeat reported the launch figures; intermediate changes were not itemized.

The current page lists no batch discounts or context-length tiers. We treat those as not offered.

3.2 Third-party hosts of the exact same model

OpenRouter's page for xiaomi/mimo-v2.5-pro routes across seven providers. Its headline price, $0.3045 in / $0.609 out with a "30% off" badge, belongs to one of them: GMICloud. It is not OpenRouter's own price and not Xiaomi's. Provider prices shown on OpenRouter on September 19, 2026:

Provider on OpenRouterInputCache hitOutputLabel
Xiaomi$0.435$0.0036$0.87(matches Xiaomi's own rate card)
AtlasCloud$0.435$0.0036$0.87THIRD-PARTY HOST PRICE
GMICloud (30% off $0.435 / $0.87)$0.3045$0.0028$0.609THIRD-PARTY HOST PRICE
NovitaAI (8% off)$0.4802$0.003956$0.9605THIRD-PARTY HOST PRICE
StreamLake$0.522$0.00432$1.044THIRD-PARTY HOST PRICE
DigitalOcean$0.48$0.096$1.80THIRD-PARTY HOST PRICE
DeepInfra (list $1.00 / $0.20 / $3.00; 61% off shown)$0.39$0.078$1.17THIRD-PARTY HOST PRICE

The 30% discount is currently displayed for GMICloud. It is a promotion and can end without notice.

3.3 Is there a real price spread? Yes, but a narrow one

With the V2.5 mix-up removed, an exact-model spread still exists in this table. Effective input prices run $0.3045–$0.522 (1.7×). Output runs $0.609–$1.80 (3.0×, DigitalOcean versus GMICloud's promo). DeepInfra's crossed-out $1.00 / $3.00 list looks like Xiaomi's old launch pricing.

Treat this as a snapshot, not an arbitrage. Promotions expire. Uptime differs: in the three-day window OpenRouter showed, GMICloud's was 88.1% versus Xiaomi's 99.4%. Host quantization, context caps and data-handling terms can also differ; verify them before choosing a provider.

3.4 Can you buy from Xiaomi directly, internationally?

Yes, per Xiaomi's documentation. The platform FAQ describes overseas top-ups by Apple Pay, Google Pay and credit or debit cards, processed through a payment gateway and settled in US dollars. Balances on the pay-as-you-go API are refundable. The rate card has an explicit overseas USD column..

The public documentation does not specify which countries are eligible, sales-tax and VAT treatment, enterprise invoicing, where overseas requests are processed, or SLAs. If your legal team needs data-residency answers, ask Xiaomi.


4. Workload costs (recalculated from scratch)

Formula (prices per 1M tokens, volumes in millions):

Cost = (Input × HitShare × P_hit) + (Input × (1 − HitShare) × P_miss) + (Output × P_out)

Baseline is Xiaomi's rate card ($0.435 / $0.0036 / $0.87). The promo column uses GMICloud's rate through OpenRouter ($0.3045 / $0.0028 / $0.609). Cache pricing is applied only in workload D, the cache-heavy case; a hit price is documented for both endpoints. The 80% hit share is a MODELING ASSUMPTION; real hit rates depend on how stable your prompt prefixes are.

WorkloadFormula (Xiaomi native)Xiaomi nativeGMICloud promo
A: 1M in + 250K out1 × 0.435 + 0.25 × 0.87$0.6525$0.4568
B: 10M in + 3M out10 × 0.435 + 3 × 0.87 = 4.35 + 2.61$6.96$4.872
C: 400M in + 100M out (monthly)400 × 0.435 + 100 × 0.87 = 174 + 87$261.00$182.70
D: 100M in + 20M out, 80% cache hits80 × 0.0036 + 20 × 0.435 + 20 × 0.87 = 0.288 + 8.70 + 17.40$26.39$18.49
D without caching, for reference100 × 0.435 + 20 × 0.87$60.90$42.63

All figures are CALCULATION. Caching cuts workload D by 57% at Xiaomi's rates. Note that cache writes are free only "for a limited time." If Xiaomi starts charging for them, cache-heavy workloads get more expensive.


5. Self-hosting economics

5.1 Free weights are not free inference

The MIT license gives you the right to run the model. It does not give you GPUs, power, networking, a serving stack or an on-call engineer. A hosted API spreads all of those across many customers.

5.2 Total parameters decide memory; active parameters decide compute

MiMo-V2.5-Pro routes each token through 8 of 384 experts, so about 42B parameters do the arithmetic for any one token. But the router can send any token to any expert, so all 1.02T parameters must sit in fast GPU memory at all times. At FP8 that is about a terabyte before you store a single token of context.

Compute scales with the 42B; the memory bill scales with the 1.02T. For small or bursty traffic, you pay for a trillion-parameter footprint to do a 42B-sized amount of work.

5.3 What the deployment documentation actually says

Xiaomi's public deployment documentation does not set a minimum hardware configuration. Available deployment recipes differ in validation, so treat them as examples rather than requirements.

SetupTotal GPU memorySourceStatus
8× H100 80GB640 GBn/aCannot hold ~1.03 TB of FP8 weights (CALCULATION)
16× H100 or H200 (2 nodes × 8, TP=16)1,280 GB / 2,256 GBSGLang cookbook; Xiaomi's model-card example commandCookbook marks it "not yet verified"
8× H200 141GB, one node, TP=81,128 GBvLLM recipe linked from Xiaomi's model cardDocumented; needs a special vLLM build
8× B200 192GB, one node, TP=81,536 GBSGLang cookbookMarked "verified"; benchmarked
TPU v7x, 16 chipsn/aSGLang (sgl-jax)Marked "verified"

The eight-H200 setup leaves very little room. At 95% memory utilization the usable pool is about 1,072 GB; subtract roughly 1,030 GB of weights and you have about 40–50 GB for context, activations and overhead (CALCULATION; a GPU-cloud vendor's guide reaches a similar figure).

A widely repeated claim that you need "16–32× H100" is not supported as a requirement. Sixteen H100s is the low end of a reported two-node recipe; public documentation does not specify a 32-GPU configuration. The 8-, 16- and 32-GPU figures below are scenarios, not vendor requirements.

5.4 Long context is a second memory bill

Xiaomi's hybrid attention (ten global layers, sixty sliding-window layers) cuts KV-cache storage by close to 7× at long context, according to the model card. It does not make it free. Using the architecture table in the model card (8 KV heads; key size 192, value size 128; 10 global layers):

Context per requestKV cache, FP8KV cache, BF16
128K tokens≈ 3.4 GB≈ 6.7 GB
1M tokens≈ 25.6 GB≈ 51.2 GB

CALCULATION: KV bytes per token = 10 layers × 8 heads × (192 + 128) × bytes per value. It ignores the small sliding-window cache, activations and fragmentation, so treat it as a floor.

On eight H200s, a single 1M-token request at FP8 KV would consume most of the 40–50 GB left after weights, which is why one GPU-cloud guide reserves 1M-context serving for 16-GPU configurations. Long-context agents on a minimal cluster mean low concurrency, so lower utilization and a higher cost per token.

5.5 The cost model

Formulas:

  • Monthly cost = GPUs × $/GPU-hour × 730 (rented, always on)
  • Owned monthly cost = capex ÷ amortization months + (kW × PUE × $/kWh × 730)
  • Capacity (M tokens/month) = tokens per second × 3,600 × 730 × utilization ÷ 1,000,000
  • Self-host $/1M = Monthly cost ÷ Capacity
  • Break-even utilization = Self-host $/1M at 100% ÷ API $/1M

Inputs and where they come from:

  • GPU prices. A GPU-cloud vendor (Spheron) listed 8-GPU nodes on July 3, 2026 at $29.58/hour on-demand and $14.08/hour spot for H200, and $42.70/hour spot for B200. That is about $3.70, $1.76 and $5.34 per GPU-hour. These are old, single-vendor numbers, so they are a MODELING ASSUMPTION, not a market quote.
  • Throughput. Two vendor-published reference points from the SGLang cookbook: (a) an 8× B200 node measured at about 4,415 tokens/second total (2,679 input + 1,736 output) on SGLang's random-prompt benchmark (nominal 1,024-token inputs and outputs, concurrency 100, no speculative decoding; the run's actual token counts were about 61% input); (b) a two-node Hopper reference (GPU model not disclosed), per node: prefill about 29,850 tokens/second at 16K context, and decode of 1,875 tokens/second without multi-token prediction (MTP) up to 5,103 with it, at a per-rank batch of 64. From (b) we derive an "agent-style" request of 16K input + 1K output taking about 0.73–1.07 node-seconds, or about 15,900–23,200 blended tokens/second. Assuming prefill and decode run back to back is a simplification. These are vendor figures, not independent measurements.
  • Owned hardware (illustrative). $300,000 per 8-GPU server over 48 months, 10.2 kW draw, PUE 1.3, $0.12/kWh. That is $6,250 + $1,162 = $7,412/month. All are MODELING ASSUMPTIONs.

Fixed monthly bill by GPU count (rented, 730 hours, linear scaling assumed):

GPUs≈$1.76/GPU-hr (spot-like)≈$3.70/GPU-hr (on-demand-like)≈$5.34/GPU-hr (B200 spot-like)
8$10,278$21,593$31,171
16$20,557$43,187$62,342
32$41,114$86,374$124,684

MODELING ASSUMPTION (rates) and CALCULATION. Doubling the GPUs doubles the bill and, if throughput scales, the capacity. So GPU count changes your minimum commitment and whether long contexts fit, not your cost per token. For scale: Workload C costs $261 on the API. The cheapest cell above is about 39× that.

Break-even utilization against Xiaomi's API (API blended price is per million tokens at the stated mix):

ScenarioNode $/hrBlended tok/sSelf-host $/1M at 100%API $/1MBreak-even utilization
Short-request benchmark (nominal 1K/1K), B200 measured$42.704,415$2.69$0.606443% (never)
Agent-style 16K/1K, no MTP, spot-like$14.0815,898$0.246$0.46153%
Agent-style, no MTP, on-demand-like$29.5815,898$0.517$0.461112% (never)
Agent-style, MTP accept-length 4, spot-like$14.0823,225$0.168$0.46137%
Agent-style, MTP accept-length 4, on-demand-like$29.5823,225$0.354$0.46177%
Same MTP case, but API gets 80% cache hits, spot-like$14.0823,225$0.168$0.136124% (never)
Same, on-demand-like$29.5823,225$0.354$0.136261% (never)
Owned server (illustrative), MTP, vs uncached API$10.1523,225$0.121$0.46126% (89% vs cached API)

CALCULATION on MODELING ASSUMPTION inputs. The API blend is 60.7% input / 39.3% output for the first row (the benchmark's measured mix), and 16:1 input:output for the agent rows. Cache-hit rows use 80% hits at $0.0036.

What the table says. Self-hosting wins on price only when the work is input-heavy, the API side gets little caching, and GPU time is cheap, and even then the node must stay roughly 25–80% utilized, depending on hardware economics and workload shape. If your traffic caches well, the API's $0.0036 hit price is hard for a self-hosted cluster to beat unless you also engineer your own prefix caching.

In volume terms, an eight-GPU node bill equals API spend on about 22B tokens a month (spot-like) to 47B (on-demand-like) in the uncached agent case. One node running flat out processes roughly 42–61B there. With heavy API caching the crossover is above 75B, beyond one node's capacity. These are not small-team volumes.

The idealized exception: renting by the minute. If you could pack a batch perfectly onto rented GPUs, Workload C (500M tokens) needs about 6–9 node-hours: roughly $84–$123 on spot-like pricing and $177–$258 on-demand-like, against the API's $261. That is parity at best, and it assumes vendor throughput, available spot capacity and perfect packing. It also ignores loading about 1 TB of weights each time you spin up (the vendor guide budgets 15–30 minutes for the download) and a 20-minute minimum billing period. With that overhead, the small Workload B (13M tokens, 0.16–0.23 node-hours of pure compute) lands at roughly $6–$22 self-hosted versus $6.96 on the API.

5.6 What these numbers leave out

The model flatters self-hosting. It excludes engineering time, redundancy (one node is one point of failure), storage and egress, spike headroom, upgrade work as inference engines change (the vLLM recipe needed a special build at launch), and security and compliance overhead. Each pushes break-even utilization up.

5.7 Why a low API price can be rational even when weights are free

A hosting provider reaches the high-utilization corner of the table above by pooling many customers' traffic onto one expensive memory footprint. You would have to get there with one company's traffic. A secondary report of Xiaomi's May 27 price announcement says the cuts came partly from inference-infrastructure gains such as better caching. And on Xiaomi's own ClawEval benchmark it claims roughly 40–60% fewer tokens per agent trajectory than some frontier peers, a vendor claim to test on your own work.

When self-hosting still makes sense:

  1. Data or code cannot leave your infrastructure.
  2. You need to fine-tune, modify or pin the exact weights. The MIT license allows it.
  3. You run sustained, input-heavy, low-cache pipelines large enough to keep nodes busy, ideally on owned or reserved hardware.
  4. You need capacity that does not depend on one vendor's pricing decisions. Xiaomi's list price has already fallen by more than half since April, which is good for buyers but shows how much sits outside your control.

6. Build & Earn: a developer product on MiMo-V2.5-Pro

Scenario (all MODELING ASSUMPTION): an AI pull-request review assistant sold to small dev teams for $29 per customer per month. Each customer triggers 300 review runs a month. A run uses 40K input tokens and 3K output tokens, so 12.9M tokens per customer per month. Model cost is Xiaomi's rate card.

CaseModel cost per customerGross margin at $29
Base usage, no caching$6.0079.3%
Base usage, 75% cache hits$2.1292.7%
3× usage (agent loops), no caching$18.0137.9%
3× usage, 75% cache hits$6.3678.1%
10× usage, no caching$60.03−107%
10× usage, 75% cache hits$21.2026.9%

CALCULATION. Formula per run: 40,000 × (hit × $0.0036 + (1 − hit) × $0.435) ÷ 1,000,000 + 3,000 × $0.87 ÷ 1,000,000, multiplied by 300 runs. At GMICloud's promo rate, the base cases are $4.20 and $1.49.

What to watch. The margins are "gross" from model cost only. They exclude payment fees, hosting, support and the web search plugin ($5 per 1,000 calls). Xiaomi's public API pricing does not clearly state how reasoning tokens are metered; if thinking output counts as billable output, your output figures could be much higher. The pattern to take away: usage variance matters more than the per-token discount. The gap between cached and uncached is big, and a 10× agent-loop customer can flip a healthy margin negative.

Would this product ever justify self-hosting? At uncached usage, the API bill equals an eight-GPU spot-like node ($10,278/month) at about 1,700 customers. A node running flat out holds roughly 3,200 customers of this shape. So you would need the node more than half full just to break even. With 75% caching, you would need about 4,800 customers, more than one node can serve. Build on the API first; revisit when you have a customer base large enough to fill dedicated hardware.