AWS Bedrock vs Azure AI Foundry vs Google Vertex AI: Real Managed LLM Platform Cost per 1 Billion Tokens in 2026
The short answer: the "cloud markup" story is not one number — it depends entirely on whether the model is proprietary (OpenAI, Anthropic) or open-weight. For Claude, there is effectively no markup at all: Anthropic's own documentation describes Google Cloud as a "partner-operated" platform pricing to match Anthropic's direct API dollar for dollar, and AWS Bedrock does the same for Claude — at 1 billion tokens (5:1 ratio), Claude Sonnet 5 costs $5,000 whether accessed through Bedrock, Vertex AI, or Anthropic's own API directly. For open-weight models, the managed-platform tax is real and large: Llama 3.3 70B costs $880 per billion tokens on Together AI, but the identical weights cost $1,364 on Google Vertex AI (a 55% premium) and $2,649 on AWS Bedrock (a roughly 201% premium) for running the exact same model. Layered on top of any of these token rates, independent cost analyses report that real production bills on AWS Bedrock specifically run 1.5 to 2 times the initial token-based estimate once Knowledge Bases, Agents, and Guardrails are added.
Claude: no markup, but a recent repricing
Claude on Bedrock and Vertex AI (OFFICIAL, cross-referenced against Anthropic's own pricing documentation). Anthropic's documentation explicitly states that Google Cloud is a partner-operated platform with independent regional pricing that, in practice, tracks Anthropic's direct API rate dollar for dollar; AWS Bedrock does the same for Claude, described directly on AWS's own Bedrock pricing page. Claude Sonnet 5 was priced at $2 input / $10 output per million tokens as an introductory rate through August 31, 2026, reverting to $3 input / $15 output afterward — identically on Bedrock, Vertex AI, and Anthropic's direct API. As of late September 2026, the current rate on all three surfaces should be the $3/$15 post-introductory rate; if a quote or calculator still shows $2/$10, it is referencing the expired introductory period. Claude Opus 4.8 runs $5/$25 per million on all three platforms identically. The "Bedrock markup" or "Vertex markup" framing simply does not apply at the token-rate level for Claude — the managed platforms are reselling Anthropic's own pricing, not adding a margin on top of it.
Open-weight models: the real, model-specific markup
Llama 3.3 70B (REPORTED, cross-referenced against Together AI's published rate). The same open-weight model, run through three different managed layers: $0.88 per million tokens on Together AI (a dedicated inference provider), $1.36 per million on Google Vertex AI Model Garden (a 55% premium for identical weights), and roughly $2.65 per million on AWS Bedrock (an approximately 201% premium, more than three times Vertex's markup on the same model). Nothing about the model's capability changes between these three surfaces — the premium buys managed infrastructure, billing integration, and access to a single API surface spanning 200-plus models on Vertex or a comparable catalog on Bedrock, not anything specific to how well the model runs.
Azure AI Foundry's small-model tier (OFFICIAL, Microsoft's Foundry pricing, mid-2026). Phi-4-mini: approximately $0.07 input / $0.23 output per million tokens, described as a 35–40x cost reduction versus GPT-4o for the same token volume — the cheapest tier in this entire comparison by a wide margin, though at a real capability tradeoff versus a frontier model. Azure's GPT-4.1-class rate is reported at $2.00 input / $8.00 output per million tokens, with a 1-million-token context window relevant for long-document or extended-reasoning workloads.
Cost per 1 billion tokens (5:1 input:output ratio)
Formula: cost per 1B = (5 × input rate + output rate) ÷ 6 × 1,000.
| Model tier | Platform | Cost per 1B tokens |
|---|---|---|
| Claude Sonnet 5 (current rate) | Bedrock, Vertex AI, or Anthropic direct — identical | $5,000 |
| Azure GPT-4.1-class | Azure AI Foundry | $3,000 |
| Llama 3.3 70B | AWS Bedrock | $2,649 |
| Llama 3.3 70B | Google Vertex AI | $1,364 |
| Llama 3.3 70B | Together AI (dedicated inference, for reference) | $880 |
| Azure Phi-4-mini | Azure AI Foundry | $97 |
The absolute dollar gap between a frontier-tier model and Phi-4-mini widens, not narrows, as usage grows, since both scale linearly with volume on their respective per-token rates.
What the token rate leaves out
Independent FinOps analysis of AWS Bedrock specifically found that most teams' actual production bills run 1.5 to 2 times their initial token-based estimate, not because of hidden fees but because the real cost of running Bedrock in production is spread across several adjacent, separately-billed services: Knowledge Bases (which depend on a separate vector-database service like OpenSearch Serverless, reported with a floor around $350/month for a minimal deployment), Agents (whose multi-step tool-calling behavior can consume 5–10 times the token volume of a single direct model call), and Guardrails (billed per text unit screened, covered in this site's companion AI-safety cost comparison). Azure's own documentation similarly notes that its pricing calculator "covers the token portion accurately but leaves out Azure AI Search, fine-tuned model hosting, monitoring infrastructure, and support costs" that push real deployment bills 15% to 40% above the token estimate. Vertex AI carries an equivalent pattern through its own adjacent services (Vector Search, Agent Builder, Document AI), though a single consolidated "real-world multiplier" figure for Vertex specifically was not found in the sources reviewed.
Sensitivity
- Model type — proprietary versus open-weight. The single largest structural fact in this article: Claude, specifically, carries no token-rate markup across Anthropic direct, Bedrock and Vertex AI; open-weight models in this comparison carry a markup of 55% to over 200% depending on cloud and model. This article does not test every proprietary model or provider, so it does not generalize the no-markup finding beyond Claude on these three surfaces.
- Which cloud hosts a given open-weight model. A roughly 3.6x difference in markup magnitude between Vertex (55%) and Bedrock (201%) on the identical Llama 3.3 70B weights.
- Adjacent-service usage. Knowledge Bases, Agents, and Guardrails can each add substantial, separately-billed cost on top of the token rate, with AWS's own ecosystem reported at a 1.5–2x real-world multiplier.
- Which Claude introductory-versus-standard rate applies. The $2/$10 introductory rate expired August 31, 2026; any quote still citing it is stale by the time this article was researched.
Budgeting traps
- Assuming a flat "cloud markup" percentage applies to every model. It does not; Claude carries none on these three surfaces, and open-weight markups vary by a factor of 3.6x between clouds on the same model.
- Budgeting Bedrock, Azure, or Vertex off the token-rate calculator alone. Real production bills on Bedrock specifically were reported to run 1.5–2x that estimate once adjacent services are included.
- Choosing a managed cloud for an open-weight model purely for convenience without checking the dedicated-inference-provider price first. The premium for the same weights can exceed 200% on some clouds.
- Using a stale Claude rate in a budget model. The $2/$10 introductory pricing on Sonnet 5 ended August 31, 2026.
What to ask before you buy
For any open-weight model, get a direct quote from a dedicated inference provider (Together AI, Fireworks, or similar) alongside your cloud platform's managed rate, since the markup for identical weights varies enormously by cloud and can exceed 200%. Ask your cloud platform for a fully-loaded cost estimate including Knowledge Bases, Agents, and Guardrails if your architecture uses them, rather than budgeting off the token-rate calculator alone.