NotCheapMAKE EVERY CREDIT COUNT

Cohere vs Mistral vs OpenAI: Real Enterprise RAG Cost per 1 Million Queries in 2026

The short answer: in a retrieval-augmented generation pipeline, the generation model is where the money goes, embeddings are close to free, and reranking becomes the second-largest line once the model is cheap. With every vendor given the same context and output lengths, an enterprise knowledge-search query (6,000 input tokens, 400 output tokens, one rerank) costs about $21,600 per million queries on Cohere's Command A, $18,550 on OpenAI's GPT-6 Sol and $6,150 on Mistral Large 3, at 1 million queries a month with an external vector database. Embeddings are under 0.1% of that. The spread comes from one column of the price sheet: token rates. Doubling the retrieved context raises the generation cost by 75–86%, while doubling the output length raises it by 14–25%.

What each vendor bills

Cohere (REPORTED, from Cohere-linked pricing summaries, June 2026). Command A at $2.50 input and $10.00 output per million tokens, Command R at $0.15 and $0.60, Command R7B at $0.0375 and $0.15. Embed 4 at $0.12 per million text tokens ($0.47 for images), Rerank 4 Fast at $2.00 per 1,000 searches and Rerank 4 Pro at $2.50 per 1,000 searches, billed per search, not per token. Cohere's documentation confirms that generation is billed per token with different input and output rates, rerank per search and embeddings per token. Cohere's enterprise products, North and Compass (enterprise search), are custom-priced (QUOTE-ONLY / UNKNOWN), and no Command A+ rate was captured.

Mistral (REPORTED). Mistral Large 3 at $0.50 input and $1.50 output per million tokens, Medium 3.5 at $1.50 and $7.50, Small 4 at $0.15 and $0.60. One source lists Large 3 at $2.00 and $5.00, so the disagreement is kept and both are modeled. Batch processing carries a 50% discount on most models (REPORTED). A cached-input rate was not listed for the models reviewed. Mistral Embed's per-token rate and any native reranking endpoint were not captured (QUOTE-ONLY / UNKNOWN), so Cohere's Embed 4 rate stands in for Mistral embeddings and a separate Cohere rerank is used, both ILLUSTRATIVE.

OpenAI (OFFICIAL). text-embedding-3-small at $0.02 per million tokens and text-embedding-3-large at $0.13 (OpenAI's model pages). For generation, this article uses GPT-6 Sol at Standard short-context pricing: $2.00 per million input tokens, $0.20 cached input and $10.00 output, with Batch pricing at $1.00 input, $0.10 cached input and $5.00 output. GPT-6 Luna, a cost-sensitive alternative, is priced at $0.10 input and $0.50 output per million tokens on Standard, which would take shape B generation to about $800 per million queries; this article prices the OpenAI stack on GPT-6 Sol and does not compare model quality.

Normalized assumptions

Every stack gets the same query shapes, so differences reflect prices, not context choices. All are ILLUSTRATIVE.

ShapeInput tokens (prompt + retrieved context)Output tokensQuery embeddingRerank
A: lightweight FAQ1,50015030 tokensnone
B: enterprise knowledge search6,00040030 tokens1 search
C: heavy document-analysis RAG30,0001,20030 tokens1 search

Corpus: 1 million chunks of 512 tokens (512 million tokens), embedded once. That is $61.44 on Cohere Embed 4, $10.24 on text-embedding-3-small and $66.56 on text-embedding-3-large, and a similar amount for monthly refreshes. Amortized over 12 months it is about $5 a month or less.

Vendor-native vector and search products (Cohere Compass, OpenAI's hosted file search and vector stores, Mistral's document libraries) either had no dependable public price or were not captured, so the model uses an external vector database as an ILLUSTRATIVE fixed cost: one node at 730 hours of a $0.75 node-hour (about $547.50 a month) up to 1 million queries a month, and four nodes ($2,200) at 10 million. Rerank is priced at Cohere's $2.00 per 1,000 searches for all stacks.

Formulas: generation = (input × input rate + output × output rate) ÷ 1,000,000 per query; embedding = 30 × rate ÷ 1,000,000; rerank = $0.002 per query where used; vector = fixed monthly cost ÷ (queries per month ÷ 1,000,000).

Cost per 1 million queries, by shape

Generation, embedding and rerank only (before the vector database):

ShapeCohere Command AMistral Large 3 ($0.50 / $1.50)Mistral Large 3 ($2.00 / $5.00)OpenAI GPT-6 Sol
A: FAQ$5,254$979$3,754$4,501
B: knowledge search$21,004$5,604$16,004$18,001
C: heavy document analysis$89,004$18,804$68,004$74,001

Including the external vector database at 1 million queries a month (about $548 plus $5 of ingestion per million queries):

ShapeCohere Command AMistral Large 3 ($0.50 / $1.50)OpenAI GPT-6 Sol
A$5,806$1,531$5,053
B$21,556$6,156$18,553
C$89,556$19,356$74,553

Effect of volume

For shape B, the vector layer is a fixed cost, so it dominates only at low volume:

Queries per monthCohere Command AMistral Large 3 ($0.50 / $1.50)OpenAI GPT-6 Sol
100,000$26,530$11,130$23,527
1 million$21,556$6,156$18,553
10 million$21,224$5,824$18,221

At 100,000 queries a month, the vector database alone is $5,475 per million queries, which is 49% of the Mistral cost. At 10 million, it is $220.

Contribution by component

Shape B at 1 million queries a month:

ComponentCohereMistral Large 3 ($0.50 / $1.50)OpenAI GPT-6 Sol
Generation88.2%58.5%86.3%
Rerank9.3%32.5%10.8%
Vector database2.5%8.9%3.0%
Embeddings0.02%0.06%0.00%

Embeddings are close to irrelevant on the query side. Rerank, which adds a flat $2,000 per million queries, is a third of the cost on the cheapest generator and about a tenth on Cohere and OpenAI.

Sensitivity to context and output length

Shape B, generation cost per million queries:

ChangeCohereMistral Large 3 ($0.50 / $1.50)OpenAI GPT-6 Sol
Baseline (6,000 in / 400 out)$19,000$3,600$16,000
Context doubled to 12,000$34,000 (+79%)$6,600 (+83%)$28,000 (+75%)
Output doubled to 800$23,000 (+21%)$4,200 (+17%)$20,000 (+25%)

Input length matters most on stacks with cheap output; output length matters most where the output rate is high relative to input. Caching and batch change the picture further: if half of the input is a cacheable static prefix, shape B generation falls by about 34% on OpenAI at its listed $0.20 cached rate ($16,000 to $10,600 per million queries: 3,000 uncached input tokens at $2, 3,000 cached at $0.20 and 400 output tokens at $10, per million tokens) and by about 36% on Cohere if a cached rate of 10% of the input price applied (an ILLUSTRATIVE assumption, since none was captured). Batch pricing applies only to asynchronous workloads: OpenAI's Batch rates ($1.00 input, $5.00 output) take shape B generation to $8,000 per million queries, and a 50% discount reported for most Mistral models takes Mistral's to $1,800.

Budgeting traps

  • Comparing headline token prices without normalizing context. A cheap model with a bloated context can cost more than a pricier model with a tight one.
  • Reranking as a flat per-search fee. At $2 per 1,000 searches it is invisible in a pilot and large next to a cheap generator.
  • Vendor-native retrieval. Hosted vector stores and enterprise search products were not comparable on public prices, so the vector line here is an external database cost.
  • Model choice hidden in "latest" aliases. Aliases can change the price without a code change.

What to ask before you buy

Send the same 100 real queries through each stack, count billed input and output tokens (Cohere reports billed tokens separately from raw tokens), and price them with the current rate card. Ask each vendor for the reranking and embedding rates in writing, and whether cached input and batch pricing apply to your model.