NotCheapMAKE EVERY CREDIT COUNT

Snowflake Cortex AI vs Databricks Mosaic AI vs Google Vertex AI: Real Enterprise GenAI Cost per 1 Million Requests in 2026

The short answer: comparing model token prices alone tells you almost nothing about the bill. Once search, embeddings, storage and runtime are added, the ranking of the three platforms changes with volume. At 100,000 requests a month, fixed platform costs make up roughly 43–78% of the total. At 10 million requests a month, model tokens make up close to 90% and the platforms converge. Cost per 1 million requests in this model ranges from roughly $1,900 to $7,500 with a low-cost model, and rises above $6,800 on Snowflake once a mid-tier model is selected.

The workload

The assumptions are ILLUSTRATIVE and identical across platforms.

  • One retrieval-augmented request: 3,000 input tokens (prompt plus retrieved context), 300 output tokens, one 30-token query embedding.
  • A corpus of 2 million chunks, 500 tokens each, 768-dimension vectors, about 1,000 bytes of text per row.
  • A general runtime allowance of $0.20 per 1,000 requests for orchestration, applied equally to all three so that structural differences show.
  • Volumes: 100,000, 1 million and 10 million requests per month.

What each platform bills

Snowflake Cortex (OFFICIAL, Snowflake documentation). AI Functions and the REST API are billed in AI Credits per million tokens, at rates that vary by model. Cortex Search bills serving compute in credits per GB per month of indexed data (including vector embeddings), running continuously while the service is resumed, plus embedding compute per token when data is inserted or updated. Snowflake's documented example uses 6.3 credits per GB per month and 0.05 credits per million embedding tokens. Warehouse compute for any SQL that runs alongside is billed separately. The dollar value of a credit depends on edition, region and contract; published examples used both $2 and $3 (REPORTED). Per-model rates were reported as ranging from about $0.12 to $5.10 per million tokens (REPORTED).

Databricks Mosaic AI (OFFICIAL, Databricks pricing pages). Everything is metered in Databricks Units (DBUs). Vector Search was listed at 4.0 DBU per hour per unit, or $0.28 per hour per unit at $0.07 per DBU in US East, with one unit holding up to 2 million 768-dimension vectors. The page's own example shows 10 million vectors on 5 units for 720 hours costing $1,008 a month. Foundation Model APIs are billed in DBU per million tokens; a reported Llama 4 Maverick rate of 7.143 DBU input and 21.429 DBU output equals about $0.50 and $1.50 per million tokens (REPORTED). A reported embedding model rate of 1.429 DBU per million tokens equals about $0.10 (REPORTED). Cloud compute is billed separately by the cloud provider in classic deployments.

Google Vertex AI (OFFICIAL, Google Cloud pricing page). Gemini 2.5 Flash standard is priced at $0.30 per million input tokens and $2.50 per million text output tokens. Vector Search index builds and updates are $3.00 per GiB. Vector Search serving cost is replicas × shards × hourly node rate × hours, and an e2-standard-16 node in us-central1 is approximately $0.7504672 per node hour, rounded here to $0.75, or about $547.50 a month for one node. The choice of an e2-standard-16 machine and a one-node deployment is ILLUSTRATIVE, because actual shard and node sizing depends on data size, queries per second, recall and latency requirements. Embedding cost is not included in the Vertex totals, which therefore slightly understate the true Vertex cost.

Platform layer, per month

Monthly requestsSnowflake ($2/credit)DatabricksVertex (one node)
100,000$122.91$221.90$584.67
1 million$305.61$404.60$764.67
10 million$2,132.61$2,231.61$2,564.67

Components: Snowflake = 51.3 credits of search serving (2 million rows × 4,072 bytes = 8.144 GB × 6.3 credits) plus query embeddings plus runtime; Databricks = one Vector Search unit at $201.60 a month plus query embeddings plus runtime; Vertex = one illustrative node at $547.50 plus a monthly index rebuild of about $17 (5.72 GiB × $3) plus runtime. At $3 per Snowflake credit, the fixed serving cost rises from about $103 to about $154 a month.

Model layer, per month

Monthly requestsVertex (Gemini 2.5 Flash)Databricks (Llama 4 Maverick)Snowflake at $0.50/MSnowflake at $2.00/MSnowflake at $5.10/M
100,000$165$195$165$660$1,683
1 million$1,650$1,950$1,650$6,600$16,830
10 million$16,500$19,500$16,500$66,000$168,300

Snowflake's model rate is a parameter because it depends on which model is selected; the three columns bracket a small open model, a mid-tier model and a frontier model on a blended per-million-token basis (3,300 tokens per request).

Total cost per 1 million requests

Monthly volumeVertexDatabricksSnowflake at $0.50/MSnowflake at $2.00/M
100,000$7,497$4,169$2,879$7,829
1 million$2,415$2,355$1,956$6,906
10 million$1,906$2,173$1,863$6,813

Formula: (platform layer + model layer) ÷ monthly requests × 1,000,000.

What the numbers show

1. Fixed serving cost dominates small workloads. At 100,000 requests a month, the platform layer is about 78% of the Vertex total, 53% of Databricks and 43% of Snowflake at $0.50 per million tokens. An always-on index charges the same whether it answers one query or a million. 2. Model choice dominates large workloads. At 10 million requests a month, the platform layer falls to about 10–13% of the total on each platform. Moving from a small to a mid-tier model on Snowflake multiplies the total bill by about 2.7x at 100,000 requests and 3.7x at 10 million. 3. A per-token winner does not survive contact with the platform layer. The cheapest model tokens can still lose at low volume because of the fixed serving line.

Sensitivity

  • Replicas. A second Vertex serving node (an ILLUSTRATIVE replica) adds $547.50 a month, which nearly doubles (+94%) the Vertex platform layer at 100,000 requests.
  • Tokens per request. Doubling retrieved context from 3,000 to 6,000 input tokens roughly doubles the model layer.
  • Runtime allowance. The $0.20 per 1,000 requests allowance adds $2,000 a month at 10 million requests; if real orchestration costs five times more, it exceeds the search layer entirely.
  • Credit price. A $3 Snowflake credit instead of $2 raises the credit-priced lines by 50%.

Budgeting traps

  • Idle indexes still bill. Search serving runs continuously while resumed, and Vertex nodes bill until undeployed.
  • Embeddings are billed twice. Once when data is indexed and again for every query.
  • Data transfer and cloud egress are not included in any total here and apply when data moves between clouds or regions.
  • Required platform consumption. A GenAI workload usually sits on top of a warehouse or lakehouse you already pay for, and that base spend is not counted here.

What to ask before you buy

Request a total-cost model that names the model, the index size, the embedding model and the expected queries per second. Ask each vendor for the minimum monthly commitment, the idle-index charge, and the rate for the exact model you plan to use.