NotCheapMAKE EVERY CREDIT COUNT

Reka vs Cohere vs Mistral: Real Enterprise LLM Cost in 2026

The short answer: on public per-token rates, enterprise LLM cost spans a factor of about 17 between the cheapest and the priciest option at the same token volume. At a 5:1 input-to-output ratio, one billion total tokens costs about $225 on Cohere's Command R, $667 on Mistral Large 3, $1,000 on Reka Flash, $2,667 on Reka Core and $3,750 on Cohere's Command A. Annual cost at 1 billion input plus 200 million output tokens a month ranges from $3,240 to $54,000. The public rate card covers only the API. Dedicated capacity, regional hosting, enterprise support and minimum commitments were not publicly priced for any of the three, with one exception: Reka states that on-premise and virtual-private-cloud licenses start at $10,000 a year.

What each vendor publishes

Reka (OFFICIAL, Reka documentation and API page). Pay-as-you-go text pricing per million tokens: Reka Edge $0.10 input and $0.10 output, Reka Flash $0.80 and $2.00, Reka Core $2.00 and $6.00, plus per-image, per-video-minute and per-audio-minute charges ($0.01–$0.02 per image, $0.06–$0.08 per video minute for Flash and Core). Reka Research is billed per 1,000 requests, at $25 standard, $35 for low parallel thinking and $60 for high parallel thinking. Reka Vision bills video indexing at $0.05 per input minute. On-premise, on-device or virtual-private-cloud deployment is a bespoke license, "as low as US$10,000 per year." Third-party pages still list older Reka Flash rates of $0.20 and $0.80 and a Reka Core rate of $2.00 and $2.00, so the current documentation figures are used. Enterprise tiers with bulk discounts are quote-only (QUOTE-ONLY / UNKNOWN).

Cohere (REPORTED). Command A at $2.50 input and $10.00 output, Command R at $0.15 and $0.60, Command R7B at $0.0375 and $0.15, all per million tokens. Embed 4 is $0.12 per million tokens and rerank is priced per 1,000 searches ($2.00 to $2.50). Older Command models are existing-customer pricing only. A second page lists Command R at $3.00 and $15.00, which is the retired April 2024 rate. Private deployment, North and Compass are quoted by Cohere sales (QUOTE-ONLY / UNKNOWN), and fine-tuning prices were not captured.

Mistral (REPORTED). Mistral Large 3 at $0.50 input and $1.50 output, Medium 3.5 at $1.50 and $7.50, Small 4 at $0.15 and $0.60 (or $0.10 and $0.30 on one page), Ministral tiers from $0.10, Magistral Medium at $2.00 and $5.00. One source lists Large 3 at $2.00 and $5.00, so both are modeled. Batch processing carries a 50% discount on most models. Cached-input pricing was not listed for the models reviewed. Enterprise, dedicated and regional-hosting terms, fine-tuning prices and minimum commitments were not publicly priced (QUOTE-ONLY / UNKNOWN).

Volume scenarios

Monthly token volumes are 100 million input plus 20 million output, 1 billion plus 200 million, and 10 billion plus 2 billion, a 5:1 ratio (ILLUSTRATIVE). Formulas: monthly cost = input millions × input rate + output millions × output rate; annual = 12 × monthly; cost per 1 billion total tokens = (5 × input rate + output rate) ÷ 6 × 1,000.

ModelCost per 1B total tokens100M in / 20M out, annual1B in / 200M out, annual10B in / 2B out, annual
Cohere Command R ($0.15 / $0.60)$225$324$3,240$32,400
Mistral Large 3 ($0.50 / $1.50)$667$960$9,600$96,000
Reka Flash ($0.80 / $2.00)$1,000$1,440$14,400$144,000
Mistral Medium 3.5 ($1.50 / $7.50)$2,500$3,600$36,000$360,000
Mistral Large 3, alternate rate ($2.00 / $5.00)$2,500$3,600$36,000$360,000
Reka Core ($2.00 / $6.00)$2,667$3,840$38,400$384,000
Cohere Command A ($2.50 / $10.00)$3,750$5,400$54,000$540,000

Enterprise contracts add support, regional hosting and commitments on top of these rates, and none of those charges were public. The table therefore shows API list cost, not a full enterprise bill. Different model tiers are not equal in capability, so the table compares price per token, not price per outcome.

Sensitivity to the input-to-output ratio

Cost per 1 billion total tokens at three ratios (input tokens per output token):

Model1:1 (generation-heavy)5:1 (baseline)20:1 (retrieval-heavy)
Cohere Command R$375$225$171
Mistral Large 3$1,000$667$548
Reka Flash$1,400$1,000$857
Reka Core$4,000$2,667$2,190
Mistral Medium 3.5$4,500$2,500$1,786
Cohere Command A$6,250$3,750$2,857

Output-heavy workloads cost more on every model. Moving from 20:1 to 1:1 multiplies cost per token by 1.6 on Reka Flash, 1.8 on Reka Core, 2.2 on Cohere Command A and Command R, and 2.5 on Mistral Medium 3.5. Models with the largest gap between output and input rates are the most sensitive to the ratio.

Batch and caching

A 50% batch discount applies to most Mistral models. On Mistral Large 3 at the baseline volumes, that lowers annual cost to $480, $4,800 and $48,000. It applies only to asynchronous jobs and cannot be combined with latency-sensitive traffic. Reka and Cohere batch discounts and cached-input rates were not captured (QUOTE-ONLY / UNKNOWN).

Private deployment break-even

Reka's on-premise license starts at $10,000 a year; the infrastructure to run it is extra. Suppose one H100 rented at $2.99 an hour costs $26,192 a year (ILLUSTRATIVE, not a statement about capacity). Total fixed cost would be about $36,192 a year, or $3,016 a month. That equals API spend on Reka Flash at 3.0 billion total tokens a month, Reka Core at 1.1 billion, Mistral Large 3 at 4.5 billion and Command A at 0.8 billion. Below those volumes, the API is cheaper on list price. Actual capacity, licensing above the floor and staffing for a private deployment were not publicly priced, so this is a screening calculation only.

Budgeting traps

  • Comparing the cheapest model of one vendor with the flagship of another. Command R at $225 per billion tokens and Command A at $3,750 are both Cohere.
  • Stale rates. Reka Flash and Cohere Command R have older listed prices on third-party pages that differ from current documentation.
  • Multimodal meters. Reka bills images, video minutes and audio minutes separately from tokens.
  • Per-request research pricing. Reka Research at $25 to $60 per 1,000 requests is a different meter from tokens.
  • Missing enterprise terms. Minimum commitments, regional hosting and support levels were not public.

What to ask before you buy

Ask each vendor for its committed-use discount, the price of dedicated capacity or a private license, regional hosting terms and the support tier that comes with your volume, in writing. Then rerun the table above with the negotiated rate and your real input-to-output ratio.