NotCheapMAKE EVERY CREDIT COUNT

AWS Bedrock vs Azure AI Foundry vs Google Vertex AI: Real Managed LLM Platform Cost per 1 Billion Tokens in 2026

The short answer: the "cloud markup" story is not one number — it depends entirely on whether the model is proprietary (OpenAI, Anthropic) or open-weight. For Claude, there is effectively no markup at all: Anthropic's own documentation describes Google Cloud as a "partner-operated" platform pricing to match Anthropic's direct API dollar for dollar, and AWS Bedrock does the same for Claude — at 1 billion tokens (5:1 ratio), Claude Sonnet 5 costs $5,000 whether accessed through Bedrock, Vertex AI, or Anthropic's own API directly. For open-weight models, the managed-platform tax is real and large: Llama 3.3 70B costs $880 per billion tokens on Together AI, but the identical weights cost $1,364 on Google Vertex AI (a 55% premium) and $2,649 on AWS Bedrock (a roughly 201% premium) for running the exact same model. Layered on top of any of these token rates, independent cost analyses report that real production bills on AWS Bedrock specifically run 1.5 to 2 times the initial token-based estimate once Knowledge Bases, Agents, and Guardrails are added.

Claude: no markup, but a recent repricing

Claude on Bedrock and Vertex AI (OFFICIAL, cross-referenced against Anthropic's own pricing documentation). Anthropic's documentation explicitly states that Google Cloud is a partner-operated platform with independent regional pricing that, in practice, tracks Anthropic's direct API rate dollar for dollar; AWS Bedrock does the same for Claude, described directly on AWS's own Bedrock pricing page. Claude Sonnet 5 was priced at $2 input / $10 output per million tokens as an introductory rate through August 31, 2026, reverting to $3 input / $15 output afterward — identically on Bedrock, Vertex AI, and Anthropic's direct API. As of late September 2026, the current rate on all three surfaces should be the $3/$15 post-introductory rate; if a quote or calculator still shows $2/$10, it is referencing the expired introductory period. Claude Opus 4.8 runs $5/$25 per million on all three platforms identically. The "Bedrock markup" or "Vertex markup" framing simply does not apply at the token-rate level for Claude — the managed platforms are reselling Anthropic's own pricing, not adding a margin on top of it.

Open-weight models: the real, model-specific markup

Llama 3.3 70B (REPORTED, cross-referenced against Together AI's published rate). The same open-weight model, run through three different managed layers: $0.88 per million tokens on Together AI (a dedicated inference provider), $1.36 per million on Google Vertex AI Model Garden (a 55% premium for identical weights), and roughly $2.65 per million on AWS Bedrock (an approximately 201% premium, more than three times Vertex's markup on the same model). Nothing about the model's capability changes between these three surfaces — the premium buys managed infrastructure, billing integration, and access to a single API surface spanning 200-plus models on Vertex or a comparable catalog on Bedrock, not anything specific to how well the model runs.

Azure AI Foundry's small-model tier (OFFICIAL, Microsoft's Foundry pricing, mid-2026). Phi-4-mini: approximately $0.07 input / $0.23 output per million tokens, described as a 35–40x cost reduction versus GPT-4o for the same token volume — the cheapest tier in this entire comparison by a wide margin, though at a real capability tradeoff versus a frontier model. Azure's GPT-4.1-class rate is reported at $2.00 input / $8.00 output per million tokens, with a 1-million-token context window relevant for long-document or extended-reasoning workloads.

Cost per 1 billion tokens (5:1 input:output ratio)

Formula: cost per 1B = (5 × input rate + output rate) ÷ 6 × 1,000.

Model tierPlatformCost per 1B tokens
Claude Sonnet 5 (current rate)Bedrock, Vertex AI, or Anthropic direct — identical$5,000
Azure GPT-4.1-classAzure AI Foundry$3,000
Llama 3.3 70BAWS Bedrock$2,649
Llama 3.3 70BGoogle Vertex AI$1,364
Llama 3.3 70BTogether AI (dedicated inference, for reference)$880
Azure Phi-4-miniAzure AI Foundry$97

The absolute dollar gap between a frontier-tier model and Phi-4-mini widens, not narrows, as usage grows, since both scale linearly with volume on their respective per-token rates.

What the token rate leaves out

Independent FinOps analysis of AWS Bedrock specifically found that most teams' actual production bills run 1.5 to 2 times their initial token-based estimate, not because of hidden fees but because the real cost of running Bedrock in production is spread across several adjacent, separately-billed services: Knowledge Bases (which depend on a separate vector-database service like OpenSearch Serverless, reported with a floor around $350/month for a minimal deployment), Agents (whose multi-step tool-calling behavior can consume 5–10 times the token volume of a single direct model call), and Guardrails (billed per text unit screened, covered in this site's companion AI-safety cost comparison). Azure's own documentation similarly notes that its pricing calculator "covers the token portion accurately but leaves out Azure AI Search, fine-tuned model hosting, monitoring infrastructure, and support costs" that push real deployment bills 15% to 40% above the token estimate. Vertex AI carries an equivalent pattern through its own adjacent services (Vector Search, Agent Builder, Document AI), though a single consolidated "real-world multiplier" figure for Vertex specifically was not found in the sources reviewed.

Sensitivity

  1. Model type — proprietary versus open-weight. The single largest structural fact in this article: Claude, specifically, carries no token-rate markup across Anthropic direct, Bedrock and Vertex AI; open-weight models in this comparison carry a markup of 55% to over 200% depending on cloud and model. This article does not test every proprietary model or provider, so it does not generalize the no-markup finding beyond Claude on these three surfaces.
  2. Which cloud hosts a given open-weight model. A roughly 3.6x difference in markup magnitude between Vertex (55%) and Bedrock (201%) on the identical Llama 3.3 70B weights.
  3. Adjacent-service usage. Knowledge Bases, Agents, and Guardrails can each add substantial, separately-billed cost on top of the token rate, with AWS's own ecosystem reported at a 1.5–2x real-world multiplier.
  4. Which Claude introductory-versus-standard rate applies. The $2/$10 introductory rate expired August 31, 2026; any quote still citing it is stale by the time this article was researched.

Budgeting traps

  • Assuming a flat "cloud markup" percentage applies to every model. It does not; Claude carries none on these three surfaces, and open-weight markups vary by a factor of 3.6x between clouds on the same model.
  • Budgeting Bedrock, Azure, or Vertex off the token-rate calculator alone. Real production bills on Bedrock specifically were reported to run 1.5–2x that estimate once adjacent services are included.
  • Choosing a managed cloud for an open-weight model purely for convenience without checking the dedicated-inference-provider price first. The premium for the same weights can exceed 200% on some clouds.
  • Using a stale Claude rate in a budget model. The $2/$10 introductory pricing on Sonnet 5 ended August 31, 2026.

What to ask before you buy

For any open-weight model, get a direct quote from a dedicated inference provider (Together AI, Fireworks, or similar) alongside your cloud platform's managed rate, since the markup for identical weights varies enormously by cloud and can exceed 200%. Ask your cloud platform for a fully-loaded cost estimate including Knowledge Bases, Agents, and Guardrails if your architecture uses them, rather than budgeting off the token-rate calculator alone.