NotCheapMAKE EVERY CREDIT COUNT

OpenAI Embeddings vs Cohere Embed vs Google Gemini Embeddings: Real Embedding Cost per 1 Billion Tokens in 2026

The short answer: raw embedding API cost is the cheapest line in almost any RAG pipeline, and the token price gap between providers matters far less than the downstream storage cost the resulting vector dimensionality creates. At 1 billion tokens, OpenAI's text-embedding-3-small costs $20 ($10 on Batch), Cohere's Embed v4 costs $100, Google's gemini-embedding-001 costs $150, and OpenAI's text-embedding-3-large costs $130 ($65 on Batch). The 7.5x spread between the cheapest and priciest option here is real money at high volume, but it is typically dwarfed by what a vector database charges to store and query the resulting vectors, especially when the highest-priced models here also produce the largest vectors.

What each vendor bills

OpenAI (OFFICIAL, OpenAI pricing documentation). text-embedding-3-small: $0.02 per million tokens, Batch API $0.01 (50% off), 1,536 dimensions (configurable smaller via the dimensions parameter). text-embedding-3-large: $0.13 per million tokens, Batch $0.065, 3,072 dimensions (also configurable down to 1,024 or 256 via Matryoshka-style truncation). text-embedding-ada-002 (legacy): $0.10 per million tokens, Batch $0.05, 1,536 fixed dimensions — still supported but superseded by the -3 family for new projects. Batch processing requires submitting a job rather than a live request and returns results within a completion window, typically 24 hours, in exchange for the 50% discount.

Cohere (REPORTED). Embed v4: $0.10 per million tokens, 1,024 dimensions (fixed, unlike OpenAI's truncatable output), with a distinct input_type parameter (search_document versus search_query) that must match between indexing and querying or similarity scores become meaningless. Cohere's batch endpoint accepts large document sets per request, reducing network overhead for bulk indexing jobs, though a confirmed batch-discount percentage comparable to OpenAI's was not found in the sources reviewed (QUOTE-ONLY / UNKNOWN for an exact Cohere batch rate).

Google (OFFICIAL, ai.google.dev pricing and Google's developer blog). gemini-embedding-001: $0.15 per million input tokens, output dimension configurable from 3,072 down to 768 or lower via Matryoshka Representation Learning, supporting over 100 languages with a 2,048-token input limit. This model replaced the older text-embedding-004, which Google scheduled for deprecation on January 14, 2026; any pricing reference to text-embedding-004 still in circulation should be treated as historical. A separately reported "Gemini Embedding 2 Preview" was described in one source at $0.20 per million tokens for text input, but this article treats that figure as REPORTED and unconfirmed pending Google's own pricing page listing it as generally available.

Cost per volume

Formula: cost = tokens (millions) × price per million.

Model1M tokens100M tokens1B tokens
OpenAI text-embedding-3-small$0.02$2.00$20.00
OpenAI text-embedding-3-small, Batch$0.01$1.00$10.00
Cohere Embed v4$0.10$10.00$100.00
OpenAI text-embedding-3-large, Batch$0.065$6.50$65.00
OpenAI text-embedding-3-large$0.13$13.00$130.00
Google gemini-embedding-001$0.15$15.00$150.00

At 1 billion tokens — roughly 700 million words, or a corpus in the range of several million typical documents — the entire spread across every model and every billing mode in this table is $140, a rounding error next to most companies' cloud infrastructure budgets, but the per-million rate still matters when embedding jobs recur (re-indexing after a chunking-strategy change, continuous ingestion of new documents) rather than running once.

Dimensionality: the cost that actually compounds

Formula: downstream storage bytes = tokens' resulting vector count × output dimensions × 4 bytes (float32). The embedding API charge is a one-time cost per token; the vector database storage and query charge recurs every month the vectors are kept. OpenAI text-embedding-3-large at its default 3,072 dimensions produces a vector 3 times the size of Cohere's fixed 1,024-dimension output, and roughly twice the size of OpenAI's own text-embedding-3-small at 1,536. For a corpus of 10 million documents, that dimensionality difference was reported to translate into tens of thousands of dollars a year in vector database storage alone — an order of magnitude larger than the embedding API cost difference between the same two models. OpenAI and Google both support configurable, truncated output dimensions via Matryoshka Representation Learning (down to 256–768 on either), which is the more consequential cost lever than which vendor's per-token rate is lowest; Cohere Embed v4, in the sources reviewed for this article, is treated as fixed at 1,024 dimensions with no equivalent truncation parameter.

Batch versus real-time

OpenAI is the only vendor in this comparison with a clearly published, confirmed batch discount: 50% off list price, applying identically to both the small and large models, in exchange for asynchronous processing (results within a completion window rather than immediately). Google's and Cohere's batch processing paths exist and were referenced as available, but a confirmed discount percentage for either was not found in the sources reviewed. For a one-time bulk indexing job (initial corpus load) rather than continuous ingestion, checking each vendor's current batch terms before committing is worth the API's documentation page rather than assuming OpenAI's 50% figure transfers.

Quality-adjusted cost is not covered here

This article prices tokens, not retrieval quality. Independent evaluations report meaningful accuracy differences between these models on domain-specific text — one legal-document migration reported a 40-point top-5 recall improvement moving from a general-purpose model to a domain-tuned one — and a cheaper model that requires more aggressive reranking or a larger retrieved-context window downstream can end up costing more in total pipeline spend than a pricier model that retrieves better on the first pass. Evaluate on your own document sample before choosing by price per token alone.

Sensitivity

  1. Dimensionality choice. Truncating text-embedding-3-large from 3,072 to 1,024 dimensions cuts downstream vector storage by roughly two-thirds while keeping the same $0.13/M token API price — the single highest-leverage decision in this comparison.
  2. Batch eligibility. OpenAI's 50% batch discount only applies to workloads that can tolerate asynchronous processing; a real-time ingestion pipeline pays the full rate.
  3. Recurrence. A one-time corpus embed is a rounding error at any of these rates; continuous re-embedding (from frequent chunking-strategy changes or a fast-growing corpus) is where the per-token rate differences compound.
  4. Which Gemini model is current. text-embedding-004 is being deprecated (January 14, 2026); confirm gemini-embedding-001 is the model actually referenced in any quote or documentation you're working from.

Budgeting traps

  • Optimizing the embedding API cost while ignoring vector storage. The token price differences here are small next to what dimensionality does to a vector database bill.
  • Re-embedding an entire corpus after a minor pipeline change. A naive full re-index on every chunking-strategy tweak can burn through a budget that a one-time embed would never approach.
  • Assuming Cohere's or Google's batch discount matches OpenAI's 50%. Neither has a confirmed, published rate as clear as OpenAI's in the sources reviewed.
  • Comparing dimension-truncated and full-dimension outputs as if they cost the same downstream. The API price is per token regardless of output dimension; the storage price is not.

What to ask before you buy

Ask each vendor for their current batch-processing discount in writing before assuming a 50% reduction applies uniformly. Model your downstream vector database cost at your chosen output dimensionality before picking an embedding model on API price alone, since that is usually the larger number.