OpenRouter vs Together AI vs Fireworks AI: Real LLM Inference Cost per 1 Billion Tokens in 2026
The short answer: Together AI and Fireworks AI host the same open-weight models within a few cents of each other, and OpenRouter is not a third price at all, it is a 5.5% surcharge on top of whichever of them (or any other provider) it routes to. At a 5:1 input-to-output ratio, one billion total tokens costs about $180–$200 on a cheap open model (Llama 3.1 8B class), $880–$900 on a mid-tier model (Llama 3.3 70B class), and $2,030–$2,483 on a frontier-scale open model (DeepSeek V4 Pro class), before OpenRouter's fee. This is not a capability comparison: a cheaper model is not automatically a fair substitute for a pricier one, and this article compares dollars per token, not quality per dollar.
What each platform actually is
OpenRouter (OFFICIAL for the fee structure, REPORTED for the exact rate). OpenRouter is a router, not an independent inference host: it passes through the underlying provider's per-token rate with no markup on the tokens themselves. Its revenue comes from a fee on credit purchases, reported at 5.5% by one detailed comparison and 5% by several other sources; this article uses 5.5% as the more recently sourced figure and flags the disagreement. A separate 5% fee applies to bring-your-own-key requests above the first 1 million free tokens a month (REPORTED). The free tier is capped at 50 requests a day without credits, 1,000 a day after a $10 minimum credit purchase, and free-model variants are capped at 20 requests a minute. Unused credits can expire one year after purchase.
Together AI (OFFICIAL, Together AI pricing/model pages). Serverless per-token pricing across a large open-model catalog. Llama 3.1 8B: $0.18 input, $0.18 output. Llama 3.3 70B: $0.88 and $0.88. DeepSeek V4 Pro: $2.10 input, $4.40 output (one source), though another current tracker lists DeepSeek V4 Pro at $1.74 and $3.48 on Together, so the two figures disagree by about 21%; this article uses the higher, more recently dated figure and shows the sensitivity below. LoRA fine-tuning is available on most Llama, Mistral and Qwen sizes at roughly $8–$12 per million training tokens (REPORTED). Dedicated endpoints and GPU clusters are priced separately and were not itemized here (QUOTE-ONLY / UNKNOWN for exact dedicated rates).
Fireworks AI (OFFICIAL, Fireworks AI pricing/model pages). Llama 3.1 8B: $0.20 and $0.20. Llama 3.3 70B: $0.90 and $0.90. DeepSeek V4 Pro: $1.74 input, $3.48 output. Fireworks and Together price within a few cents of each other on most shared SKUs; Fireworks is reported as slightly cheaper on the largest Llama models ($3.00 vs $3.50 per million on Llama 3.1 405B in one comparison). Serverless is the self-serve entry point; on-demand dedicated GPU pricing (H100, H200, B200, B300 classes) and Enterprise terms are custom (QUOTE-ONLY / UNKNOWN).
Normalized volumes and models
Volumes are 100 million input plus 20 million output tokens a month, 1 billion plus 200 million, and 10 billion plus 2 billion, all at a 5:1 ratio (ILLUSTRATIVE). Three model tiers, all open-weight and hosted by both platforms: cheap (Llama 3.1 8B class), mid (Llama 3.3 70B class), and large (DeepSeek V4 Pro class).
Formula: monthly cost = input millions × input rate + output millions × output rate; OpenRouter cost = monthly cost × 1.055.
Monthly and annual cost by tier and volume
| Tier | Volume | Together AI | Fireworks AI | OpenRouter (at 5.5% over the routed provider) |
|---|---|---|---|---|
| Cheap | 100M in / 20M out | $22/mo ($259/yr) | $24/mo ($288/yr) | $23–$25/mo |
| Cheap | 1B in / 200M out | $216/mo ($2,592/yr) | $240/mo ($2,880/yr) | $228–$253/mo |
| Cheap | 10B in / 2B out | $2,160/mo ($25,920/yr) | $2,400/mo ($28,800/yr) | $2,279–$2,532/mo |
| Mid | 100M in / 20M out | $106/mo ($1,267/yr) | $108/mo ($1,296/yr) | $111–$114/mo |
| Mid | 1B in / 200M out | $1,056/mo ($12,672/yr) | $1,080/mo ($12,960/yr) | $1,114–$1,139/mo |
| Mid | 10B in / 2B out | $10,560/mo ($126,720/yr) | $10,800/mo ($129,600/yr) | $11,141–$11,394/mo |
| Large | 100M in / 20M out | $298/mo ($3,576/yr) | $244/mo ($2,923/yr) | $257–$314/mo |
| Large | 1B in / 200M out | $2,980/mo ($35,760/yr) | $2,436/mo ($29,232/yr) | $2,570–$3,144/mo |
| Large | 10B in / 2B out | $29,800/mo ($357,600/yr) | $24,360/mo ($292,320/yr) | $25,700–$31,439/mo |
At the large tier, Fireworks is cheaper than Together by about 18% in this model, and OpenRouter's fee adds 5.5% on top of whichever underlying provider it routes a given request to; OpenRouter itself is never the cheapest option, only the most convenient one.
Cost per 1 billion total tokens
Formula: cost per 1B = (5 × input rate + output rate) ÷ 6 × 1,000, at the 5:1 ratio.
| Tier | Together AI | Fireworks AI |
|---|---|---|
| Cheap (Llama 3.1 8B class) | $180 | $200 |
| Mid (Llama 3.3 70B class) | $880 | $900 |
| Large (DeepSeek V4 Pro class) | $2,483 | $2,030 |
The large tier shows the biggest platform gap: choosing Fireworks over Together on this specific model saves about $453 per billion tokens, or roughly 18%, at the higher of the two conflicting Together rates reported.
Effect of output-heavy workloads
DeepSeek V4 Pro's output rate is more than double its input rate on both platforms, so shifting the ratio changes cost more here than on the flatter-priced cheap and mid tiers.
| Ratio (input:output) | Cheap, Together | Mid, Together | Large, Together |
|---|---|---|---|
| 20:1 (retrieval-heavy) | $180 | $880 | $1,220 |
| 5:1 (baseline) | $180 | $880 | $2,483 |
| 1:1 (generation-heavy) | $180 | $880 | $3,250 |
Because Llama 3.1 8B and Llama 3.3 70B both price input and output identically on Together, the ratio does not change their cost per billion tokens at all. DeepSeek V4 Pro's cost per billion tokens moves from $1,220 to $3,250, a 2.7x range, purely from the input-output mix.
Caching and batch discounts
Neither Together nor Fireworks was documented with a general cached-input discount in the sources reviewed for these specific models (Together's DeepSeek listing shows a $0.20 cached-input rate elsewhere in its catalog, but this was not confirmed for the V4 Pro SKU used here; QUOTE-ONLY / UNKNOWN for this model). Batch or asynchronous discounts were not publicly listed for Together or Fireworks serverless inference in the sources reviewed; other providers in the broader market were reported with 30–50% batch discounts, so the absence of a public rate here is a gap, not a confirmation that no discount exists.
Sensitivity
- Which Together DeepSeek V4 Pro rate is real. At the lower reported rate ($1.74/$3.48, matching Fireworks exactly), Together's large-tier cost per billion tokens falls from $2,483 to $2,030, eliminating Fireworks' apparent advantage entirely.
- OpenRouter's fee rate. At 5% instead of 5.5%, the added cost on a $10,560 monthly mid-tier bill falls from $581 to $528, a modest but real difference at scale.
- Model substitution. Moving from the large to the mid tier cuts cost per billion tokens by roughly 65%, a bigger lever than choosing between Together and Fireworks on the same model.
Budgeting traps
- Treating OpenRouter as a fourth pricing option. It is a markup on top of Together, Fireworks or any other routed provider, not an independent price point.
- Assuming identical model names mean identical prices. DeepSeek V4 Pro's Together rate is disputed across sources by about 21%; always check the provider's own current page.
- Ignoring the output rate on reasoning-heavy models. DeepSeek V4 Pro's output price is more than double its input price, so a chat-style low-output workload is deceptively cheap to estimate from the input rate alone.
- Comparing capability-mismatched models by price alone. A cheap model and a frontier model are not interchangeable; this article prices tokens, not outcomes.
What to ask before you buy
Get the current per-token rate directly from Together's and Fireworks' own pricing pages for the exact model and quantization you plan to use, and confirm whether OpenRouter's current fee is 5% or 5.5% before estimating routed-traffic cost at scale.
ARTICLE 5