NotCheapMAKE EVERY CREDIT COUNT

Mistral Large 4 Pricing 2026: API Cost, Launch Discount and Open-Weight Economics

The short answer: Mistral Large 4 lists at $1.36 per million input tokens, $0.14 per million cached input tokens and $4.18 per million output tokens. Mistral's pricing page currently shows a temporary sale price of exactly half that: $0.68 input, $0.07 cached input, $2.09 output. The list price is the number to build a long-term budget on; the sale price is a launch bonus. Even at the sale price, Large 4 costs about 1.4 times as much per token as Mistral Large 3 ($0.50 / $1.50), so it is a step up in price tier, not a replacement at the same cost. The model is in public preview, and its open weights have been promised but are not yet downloadable.

What Mistral Large 4 costs

All prices are US dollars per million tokens, from Mistral's pricing page on 7 October 2026.

InputCached inputOutput
List price (durable baseline)$1.36$0.14$4.18
Launch sale price (temporary)$0.68$0.07$2.09

The API model ID is mistral-large-4. The model is a mixture-of-experts system with roughly 1.05 trillion total parameters, a 1M-token context window and a 1.6B-parameter vision encoder for image input. Mistral's own materials give slightly different figures for active parameters per token: the launch announcement says 49 billion, while the model documentation says 52 billion. The difference does not affect API billing, which is per token, but it is worth knowing if you compare the model with others on size.

Mistral also lists batch, priority and regional rates on separate tabs of its pricing page. This article compares standard API rates only; if you plan to use one of those modes, price it from that tab.

List price versus the launch discount

The sale price is 50% off every line. Mistral's pricing page shows it as a "Sale price" and does not state an end date. Mistral's changelog describes 50% off for two weeks, but the pricing page does not give a calendar end date or expiry time. The practical rule:

  • Budget on the list price. If the sale ends mid-quarter, a budget built on $0.68 / $2.09 doubles overnight.
  • Treat the sale price as savings, not as the plan. Whatever you spend while it lasts is a bonus.
  • Expect pricing to be reviewed. Large 4 is in public preview, and rates can change when a model reaches general availability.

API cost per workload

The three workloads below are illustrative, not forecasts. They use a 10-to-1 input-to-output ratio, which is typical of summarization, classification and retrieval-heavy assistants. Formula: (input tokens ÷ 1M × input rate) + (output tokens ÷ 1M × output rate).

WorkloadList-price arithmeticList totalSale-price arithmeticSale total
10M in + 1M out10 × $1.36 + 1 × $4.18 = $13.60 + $4.18$17.7810 × $0.68 + 1 × $2.09 = $6.80 + $2.09$8.89
100M in + 10M out100 × $1.36 + 10 × $4.18 = $136.00 + $41.80$177.80100 × $0.68 + 10 × $2.09 = $68.00 + $20.90$88.90
1B in + 100M out1,000 × $1.36 + 100 × $4.18 = $1,360 + $418$1,778.001,000 × $0.68 + 100 × $2.09 = $680 + $209$889.00

The blended price at a 10-to-1 mix is about $1.62 per million total tokens at list price ($17.78 ÷ 11) and about $0.81 at the sale price.

An output-heavy workload

Output tokens cost about three times as much as input tokens, so workloads that generate long text change the picture. Take 10M input tokens and 30M output tokens, as in drafting long reports or generating code from short prompts:

ArithmeticTotal
List10 × $1.36 + 30 × $4.18 = $13.60 + $125.40$139.00
Sale10 × $0.68 + 30 × $2.09 = $6.80 + $62.70$69.50

Here output is 90% of the list-price bill. If your product writes more than it reads, the output rate is the number that matters, and caching does nothing for it.

The effect of prompt caching

Cached input is priced at about one tenth of the standard input rate. Take 100M input tokens with 80% served from cache, plus 10M output tokens:

ArithmeticTotalVersus no caching
List20 × $1.36 + 80 × $0.14 + 10 × $4.18 = $27.20 + $11.20 + $41.80$80.20$177.80
Sale20 × $0.68 + 80 × $0.07 + 10 × $2.09 = $13.60 + $5.60 + $20.90$40.10$88.90

Caching cuts this workload's list-price bill by 55%. It works only when a large part of the prompt repeats between requests, such as a fixed system prompt or a shared document.

How Large 4 compares with earlier Mistral models

Model (standard API)InputCached inputOutput10M in + 1M out100M in + 10M out
Mistral Large 4, list$1.36$0.14$4.18$17.78$177.80
Mistral Large 4, sale$0.68$0.07$2.09$8.89$88.90
Mistral Large 3$0.50$0.05$1.50$6.50$65.00
Mistral Medium 3.5$1.50$0.15$7.50$22.50$225.00
Mistral Small 4$0.15$0.015$0.60$2.10$21.00

At list price, Large 4 costs 2.72 times as much as Large 3 on input and 2.79 times as much on output. At the sale price the multiples fall to 1.36 and 1.39. Medium 3.5 has a higher output price than Large 4 ($7.50, against $4.18 list or $2.09 sale), so for output-heavy work Large 4 is cheaper per token than Medium 3.5 even at list price. Mistral also hosts Z.ai's GLM 5.3 at $1.40 input, $0.14 cached and $4.40 output, which puts Large 4's list price within a few percent of it ($18.40 against $17.78 for 10M in + 1M out).

Comparison with Claude

The most direct price comparison is with Anthropic's mid-tier model. Prices below come from Anthropic's published pricing page.

ModelInputCached input (cache hit)Output10M in + 1M out100M in + 10M out10M in + 30M out
Mistral Large 4, list$1.36$0.14$4.18$17.78$177.80$139.00
Mistral Large 4, sale$0.68$0.07$2.09$8.89$88.90$69.50
Claude Sonnet 5.5$2.00$0.20$10.00$30.00$300.00$320.00
Claude Opus 5.5$4.00$0.20$20.00$60.00$600.00$640.00

Sonnet 5.5 arithmetic for 10M in + 1M out: 10 × $2 + 1 × $10 = $30. Opus 5.5: 10 × $4 + 1 × $20 = $60. At list price, Large 4 is about 41% cheaper than Sonnet 5.5 on the 10-to-1 workload and about 57% cheaper on the output-heavy one, and about 70% and 78% cheaper than Opus 5.5. Anthropic also publishes cache-write and batch rates that change these totals for cache-heavy or asynchronous work; see its pricing page for them.

These tables compare price, not quality. A cheaper model that needs more retries, longer prompts or human review can cost more per finished task. Test your own workload on each candidate before choosing on price.

What a 1M-token context window costs

A large context window makes long prompts possible; it does not make them cheap. Mistral lists a single rate for each token type and does not show a separate long-context tier on its pricing page, so every input token costs the same whether the prompt is 2,000 tokens or 900,000.

PromptList priceSale price
One 1,000,000-token prompt, uncached$1.36$0.68
The same prompt read from cache$0.14$0.07
4,000 tokens of output$0.01672$0.00836
1,000 requests with an 800,000-token prompt and 2,000-token reply, uncached800 × $1.36 + 2 × $4.18 = $1,088 + $8.36 = $1,096.36800 × $0.68 + 2 × $2.09 = $544 + $4.18 = $548.18

The same 1,000 requests on Claude Sonnet 5.5, which also bills 1M-token prompts at its standard rate, cost 800 × $2 + 2 × $10 = $1,620. So the economic question is not "can I afford one long prompt" (a full 1M-token prompt costs about $1.36) but "how many times will I send it". Repeated long prompts are where caching decides the bill: sending the same 800,000-token document 1,000 times costs about $1,096 uncached at list price, but reading it from cache for all but the first request brings the input side down by roughly 90%.

Mistral's pricing page can change; confirm at general availability that long prompts still bill at the standard rate.

Open weights: what is confirmed and what is not

Confirmed by Mistral: Large 4 is available through Mistral's API today as a public preview, and Mistral has said the model weights will be released by the end of October.

Not confirmed:

  • The release date. Third-party reports cite 27 October, but Mistral has said "by the end of the month," not a specific day.
  • The licence. Mistral has not named the licence. Third-party reports describe it as a custom licence, which could include commercial or revenue restrictions. Read the actual terms on release day.
  • Hardware requirements. Mistral has not published requirements for running the model. A model with roughly 1.05 trillion total parameters is a very large download, but how much memory and how many GPUs a production deployment needs depends on precision, serving software and context length, and none of that has been published.

Until the weights are released, self-hosting is not possible, and any self-hosting cost is hypothetical. For that reason this article does not offer a self-hosting break-even. A break-even needs a licence you can read, a hardware requirement you can price and a throughput figure you can measure. When those exist, the comparison is simple to build: monthly GPU and operations cost divided by monthly tokens served, set against $1.36 and $4.18 per million tokens (or $0.14 for cached input).

What open weights can change today is leverage, not cost. If the licence allows it, they give a buyer a credible alternative to the API, which matters in negotiation and for data-residency planning. They do not lower the API price you pay now.

Who should use Mistral Large 4

  • Teams that want a large-context model below Sonnet-class pricing. At list price it costs roughly 40–60% less than Claude Sonnet 5.5 on the workloads above, and the 1M window is billed at the standard rate.
  • Output-heavy products, where Large 4's $4.18 output rate is well below Medium 3.5's $7.50.
  • Teams already on Mistral Large 3 who need long context or vision, and can accept a per-token price about 2.7 times higher at list.
  • Not the cheapest option for simple tasks. Mistral Small 4 costs $2.10 for the 10M-in, 1M-out workload where Large 4 costs $17.78 at list, and Large 3 costs $6.50. If a smaller model passes your quality bar, use it.
  • Not yet the right model for production commitments that depend on stable pricing or self-hosting, since it is in preview and its weights are not released.

Budgeting traps

  1. Budgeting on the sale price. The sale price is temporary and has no stated end date. Use $1.36 / $4.18 for any plan that runs beyond the current week.
  2. Preview pricing. Rates can change at general availability.
  3. Long prompts sent repeatedly. One full 1M-token prompt is cheap, but a thousand of them are not. Build cost per request into the design.
  4. Output volume. Output is billed at about three times the input rate, and reasoning or long-form generation inflates it.
  5. Assuming caching is automatic. Cached input is billed at $0.14, but only the portion of a prompt that is actually served from cache earns that rate. Check usage reports before assuming a hit rate.
  6. Assuming self-hosting will be cheaper. There is no licence, hardware requirement or throughput figure to build that case on yet.
  7. Mixing modes. Batch, priority and regional rates are priced separately from the standard rates in this article.

FAQ

How much does Mistral Large 4 cost?

At list price, $1.36 per million input tokens, $0.14 per million cached input tokens and $4.18 per million output tokens. A temporary sale price currently shows $0.68, $0.07 and $2.09.

Is the Mistral Large 4 launch discount permanent?

No. Mistral labels it a sale price. The pricing page does not state an end date, so use the list price for planning.

How much does 100 million input tokens and 10 million output tokens cost?

$177.80 at list price ($136.00 input plus $41.80 output) and $88.90 at the sale price.

Is Mistral Large 4 cheaper than Claude Sonnet 5.5?

Per token, yes. At list price it is $1.36 / $4.18 against $2.00 / $10.00 for Sonnet 5.5. Whether it is cheaper per finished task depends on how well each model does your work.

How does Mistral Large 4 compare with Mistral Large 3?

Large 4 costs about 2.7 times as much at list ($1.36 / $4.18 against $0.50 / $1.50) and about 1.4 times as much at the sale price. It adds a 1M-token context window and vision input.

Does a 1M context window cost extra?

Mistral's pricing page lists a single rate per token type for the model, so a long prompt costs the same per token as a short one. A full 1,000,000-token prompt is $1.36 at list price, uncached.

Can I download Mistral Large 4 and run it myself?

Not yet. Mistral has said the weights will be released by the end of October. The licence and hardware requirements have not been published.

Will self-hosting be cheaper than the API?

That cannot be calculated until the licence, hardware requirements and real throughput are known.

What is the difference between the 49B and 52B active parameter figures?

Mistral's announcement says 49 billion active parameters and its model documentation says 52 billion. API pricing is per token and does not depend on either figure.

Related guides

Sources

Share