Reka vs Cohere vs Mistral: Real Enterprise LLM Cost in 2026
The short answer: on public per-token rates, enterprise LLM cost spans a factor of about 17 between the cheapest and the priciest option at the same token volume. At a 5:1 input-to-output ratio, one billion total tokens costs about $225 on Cohere's Command R, $667 on Mistral Large 3, $1,000 on Reka Flash, $2,667 on Reka Core and $3,750 on Cohere's Command A. Annual cost at 1 billion input plus 200 million output tokens a month ranges from $3,240 to $54,000. The public rate card covers only the API. Dedicated capacity, regional hosting, enterprise support and minimum commitments were not publicly priced for any of the three, with one exception: Reka states that on-premise and virtual-private-cloud licenses start at $10,000 a year.
What each vendor publishes
Reka (OFFICIAL, Reka documentation and API page). Pay-as-you-go text pricing per million tokens: Reka Edge $0.10 input and $0.10 output, Reka Flash $0.80 and $2.00, Reka Core $2.00 and $6.00, plus per-image, per-video-minute and per-audio-minute charges ($0.01–$0.02 per image, $0.06–$0.08 per video minute for Flash and Core). Reka Research is billed per 1,000 requests, at $25 standard, $35 for low parallel thinking and $60 for high parallel thinking. Reka Vision bills video indexing at $0.05 per input minute. On-premise, on-device or virtual-private-cloud deployment is a bespoke license, "as low as US$10,000 per year." Third-party pages still list older Reka Flash rates of $0.20 and $0.80 and a Reka Core rate of $2.00 and $2.00, so the current documentation figures are used. Enterprise tiers with bulk discounts are quote-only (QUOTE-ONLY / UNKNOWN).
Cohere (REPORTED). Command A at $2.50 input and $10.00 output, Command R at $0.15 and $0.60, Command R7B at $0.0375 and $0.15, all per million tokens. Embed 4 is $0.12 per million tokens and rerank is priced per 1,000 searches ($2.00 to $2.50). Older Command models are existing-customer pricing only. A second page lists Command R at $3.00 and $15.00, which is the retired April 2024 rate. Private deployment, North and Compass are quoted by Cohere sales (QUOTE-ONLY / UNKNOWN), and fine-tuning prices were not captured.
Mistral (REPORTED). Mistral Large 3 at $0.50 input and $1.50 output, Medium 3.5 at $1.50 and $7.50, Small 4 at $0.15 and $0.60 (or $0.10 and $0.30 on one page), Ministral tiers from $0.10, Magistral Medium at $2.00 and $5.00. One source lists Large 3 at $2.00 and $5.00, so both are modeled. Batch processing carries a 50% discount on most models. Cached-input pricing was not listed for the models reviewed. Enterprise, dedicated and regional-hosting terms, fine-tuning prices and minimum commitments were not publicly priced (QUOTE-ONLY / UNKNOWN).
Volume scenarios
Monthly token volumes are 100 million input plus 20 million output, 1 billion plus 200 million, and 10 billion plus 2 billion, a 5:1 ratio (ILLUSTRATIVE). Formulas: monthly cost = input millions × input rate + output millions × output rate; annual = 12 × monthly; cost per 1 billion total tokens = (5 × input rate + output rate) ÷ 6 × 1,000.
| Model | Cost per 1B total tokens | 100M in / 20M out, annual | 1B in / 200M out, annual | 10B in / 2B out, annual |
|---|---|---|---|---|
| Cohere Command R ($0.15 / $0.60) | $225 | $324 | $3,240 | $32,400 |
| Mistral Large 3 ($0.50 / $1.50) | $667 | $960 | $9,600 | $96,000 |
| Reka Flash ($0.80 / $2.00) | $1,000 | $1,440 | $14,400 | $144,000 |
| Mistral Medium 3.5 ($1.50 / $7.50) | $2,500 | $3,600 | $36,000 | $360,000 |
| Mistral Large 3, alternate rate ($2.00 / $5.00) | $2,500 | $3,600 | $36,000 | $360,000 |
| Reka Core ($2.00 / $6.00) | $2,667 | $3,840 | $38,400 | $384,000 |
| Cohere Command A ($2.50 / $10.00) | $3,750 | $5,400 | $54,000 | $540,000 |
Enterprise contracts add support, regional hosting and commitments on top of these rates, and none of those charges were public. The table therefore shows API list cost, not a full enterprise bill. Different model tiers are not equal in capability, so the table compares price per token, not price per outcome.
Sensitivity to the input-to-output ratio
Cost per 1 billion total tokens at three ratios (input tokens per output token):
| Model | 1:1 (generation-heavy) | 5:1 (baseline) | 20:1 (retrieval-heavy) |
|---|---|---|---|
| Cohere Command R | $375 | $225 | $171 |
| Mistral Large 3 | $1,000 | $667 | $548 |
| Reka Flash | $1,400 | $1,000 | $857 |
| Reka Core | $4,000 | $2,667 | $2,190 |
| Mistral Medium 3.5 | $4,500 | $2,500 | $1,786 |
| Cohere Command A | $6,250 | $3,750 | $2,857 |
Output-heavy workloads cost more on every model. Moving from 20:1 to 1:1 multiplies cost per token by 1.6 on Reka Flash, 1.8 on Reka Core, 2.2 on Cohere Command A and Command R, and 2.5 on Mistral Medium 3.5. Models with the largest gap between output and input rates are the most sensitive to the ratio.
Batch and caching
A 50% batch discount applies to most Mistral models. On Mistral Large 3 at the baseline volumes, that lowers annual cost to $480, $4,800 and $48,000. It applies only to asynchronous jobs and cannot be combined with latency-sensitive traffic. Reka and Cohere batch discounts and cached-input rates were not captured (QUOTE-ONLY / UNKNOWN).
Private deployment break-even
Reka's on-premise license starts at $10,000 a year; the infrastructure to run it is extra. Suppose one H100 rented at $2.99 an hour costs $26,192 a year (ILLUSTRATIVE, not a statement about capacity). Total fixed cost would be about $36,192 a year, or $3,016 a month. That equals API spend on Reka Flash at 3.0 billion total tokens a month, Reka Core at 1.1 billion, Mistral Large 3 at 4.5 billion and Command A at 0.8 billion. Below those volumes, the API is cheaper on list price. Actual capacity, licensing above the floor and staffing for a private deployment were not publicly priced, so this is a screening calculation only.
Budgeting traps
- Comparing the cheapest model of one vendor with the flagship of another. Command R at $225 per billion tokens and Command A at $3,750 are both Cohere.
- Stale rates. Reka Flash and Cohere Command R have older listed prices on third-party pages that differ from current documentation.
- Multimodal meters. Reka bills images, video minutes and audio minutes separately from tokens.
- Per-request research pricing. Reka Research at $25 to $60 per 1,000 requests is a different meter from tokens.
- Missing enterprise terms. Minimum commitments, regional hosting and support levels were not public.
What to ask before you buy
Ask each vendor for its committed-use discount, the price of dedicated capacity or a private license, regional hosting terms and the support tier that comes with your volume, in writing. Then rerun the table above with the negotiated rate and your real input-to-output ratio.