NotCheapMAKE EVERY CREDIT COUNT

Groq vs Cerebras vs SambaNova: Real High-Speed AI Inference Cost per 1 Billion Tokens in 2026

The short answer: all three sell speed on custom silicon (Groq's LPU, Cerebras's wafer-scale chip, SambaNova's reconfigurable dataflow unit), and speed and price move independently of each other. Cerebras is the fastest on Llama 3.3 70B, at roughly 2,100–2,500 tokens per second against Groq's independently benchmarked 294–750 and SambaNova's 430, but it is also the most expensive per token on every model in this comparison. At a 5:1 input-to-output ratio, one billion total tokens on Llama 3.3 70B costs about $623 on Groq, $700 on SambaNova, and $908 on Cerebras, a 46% spread between the cheapest and priciest, running in the opposite direction from the speed ranking. Faster inference is not free, and this article does not claim it produces a return; it only prices the speed.

What each vendor bills

Groq (OFFICIAL, Groq pricing page). Per-token pricing on a curated catalog of open-weight models, built on custom LPU (Language Processing Unit) chips. Llama 3.1 8B Instant: $0.05 input, $0.08 output. GPT-OSS 120B: $0.15 and $0.60. Llama 3.3 70B Versatile: $0.59 and $0.79. Qwen 3.6 27B: $0.60 and $3.00, a 5x output premium that is easy to miss when pricing off the input rate alone. Some larger or newer models (Minimax M2.7) are enterprise-only, priced on request. Free tier: 30 requests a minute on most models, no credit card required, but with strict production-blocking rate limits (REPORTED). Speech models (Whisper, TTS) are billed per hour or per character, not per token, and are a separate meter from text inference.

Cerebras (OFFICIAL, Cerebras pricing page and Artificial Analysis benchmarks). Runs open-source models only on its wafer-scale engine. Llama 3.1 8B: $0.10 input and output (flat). GPT-OSS 120B: $0.35 and $0.75. Qwen3 235B (A22B): $0.60 and $1.20. Llama 3.3 70B: $0.85 and $1.20 (from a separate comparison table; Cerebras's own model-pricing page did not list this specific SKU in the sources reviewed, so this figure carries a slightly lower confidence than the others). Free tier: 1 million tokens a day, a materially larger daily allowance than Groq's rate-limited free tier. A Code Pro subscription was reported at 24 million tokens a day for $50 a month, roughly 720 million tokens a month, which undercuts pay-per-token rates for a narrow, code-focused workload but locks the buyer into Cerebras's limited model catalog (REPORTED). No proprietary models; enterprise customers can reportedly bring custom weights.

SambaNova Cloud (OFFICIAL, SambaNova pricing page). Runs on Reconfigurable Dataflow Units. Llama 3.1 8B Instruct: $0.10 and $0.20. GPT-OSS 120B: $0.22 and $0.59. Llama 3.3 70B Instruct: $0.60 and $1.20. DeepSeek-V3.1: $3.00 and $4.50. DeepSeek-R1-0528: $5.00 and $7.00. A $5 free credit is offered on signup, reported to be consumed quickly at these rates. Enterprise and dedicated-capacity pricing is custom (QUOTE-ONLY / UNKNOWN).

Reported throughput, for context

Throughput figures vary by source and measurement method, and this article flags the disagreement rather than picking a winner. On Llama 3.3 70B: Cerebras's own comparison page reports Groq at 403–750 tokens/second against Cerebras's 2,100–2,500, a 3–6x gap; an independent benchmark (Artificial Analysis) measured Groq at 293.6 tokens/second on the same model, widening the reported gap to roughly 7–8x. SambaNova is reported at 430 tokens/second on the same model. On GPT-OSS 120B: Cerebras claims about 3,000 tokens/second on its own documentation, while Artificial Analysis measured a sustained ~1,700 independently; Groq is reported at 478–493.

Cost per 1 billion total tokens (5:1 input:output ratio)

Formula: cost per 1B = (5 × input rate + output rate) ÷ 6 × 1,000.

ModelGroqCerebrasSambaNova
Llama 3.1 8B$55$100$117
GPT-OSS 120B$225$417$282
Llama 3.3 70B$623$908$700

Groq is the cheapest on every model in this set; Cerebras is the most expensive on every model, and is also the fastest reported on the two models (Llama 3.3 70B and GPT-OSS 120B) where this article has throughput data for all three vendors; throughput on Llama 3.1 8B and Qwen 3.6 27B was not captured for this comparison, so the speed ranking should not be assumed to extend to every model in the pricing tables above. SambaNova sits between the two on GPT-OSS 120B and Llama 3.3 70B, but its output-heavy Llama 3.1 8B rate ($0.20 output versus Groq's $0.08) makes it the priciest of the three on that specific small model.

Volume scenarios

Formula: monthly cost = input millions × input rate + output millions × output rate, at 100 million input plus 20 million output tokens a month, 1 billion plus 200 million, and 10 billion plus 2 billion.

ModelVolumeGroqCerebrasSambaNova
Llama 3.3 70B100M/20M$75/mo ($898/yr)$109/mo ($1,308/yr)$84/mo ($1,008/yr)
Llama 3.3 70B1B/200M$748/mo ($8,976/yr)$1,090/mo ($13,080/yr)$840/mo ($10,080/yr)
Llama 3.3 70B10B/2B$7,480/mo ($89,760/yr)$10,900/mo ($130,800/yr)$8,400/mo ($100,800/yr)
GPT-OSS 120B100M/20M$27/mo ($324/yr)$50/mo ($600/yr)$34/mo ($406/yr)
GPT-OSS 120B1B/200M$270/mo ($3,240/yr)$500/mo ($6,000/yr)$338/mo ($4,056/yr)
GPT-OSS 120B10B/2B$2,700/mo ($32,400/yr)$5,000/mo ($60,000/yr)$3,380/mo ($40,560/yr)

Cost per 1 million requests, normalized token count

Using an ILLUSTRATIVE 500 input plus 500 output tokens per request (a typical short interactive exchange), on Llama 3.3 70B:

ProviderCost per requestCost per 1 million requestsReported latency for 1,000 tokens
Cerebras$0.00103$1,025~0.4–0.5 seconds
SambaNova$0.00090$900~2.3 seconds
Groq$0.00069$690~1.3–3.4 seconds (benchmark-dependent: 1.3s at the vendor-reported 750 tokens/second, 3.4s at the independently benchmarked 293.6 tokens/second)

Cerebras costs 49% more per 1 million requests than Groq while delivering roughly 3–8x the throughput, depending on which benchmark is used.

Interactive chat, bulk asynchronous, and real-time agent shapes

Three usage shapes, all ILLUSTRATIVE, differ in how much latency matters rather than in token volume: A, interactive chat (latency directly affects user experience, favors low time-to-first-token); B, bulk asynchronous processing (a batch job with no user waiting, favors the lowest per-token price regardless of speed); C, real-time agent workloads (multiple sequential model calls per task, where per-call latency compounds across a chain). For shape B, Groq's lower per-token price dominates the decision, since nothing is waiting on the response. For shapes A and C, Cerebras's throughput advantage is the reason to consider paying its premium, but only if the workload is genuinely latency-sensitive.

The value of latency reduction: an illustrative threshold only

This is a threshold calculation, not a claim that faster inference produces revenue. If Cerebras costs $335 more per 1 million requests than Groq on Llama 3.3 70B ($1,025 − $690), and each request completes roughly 1.5 seconds faster on Cerebras (ILLUSTRATIVE, within the reported throughput ranges), the incremental cost per second of latency saved is $0.000223 per request. Whether that is worth paying depends entirely on what the requesting application does with the saved time: a synchronous voice agent that cannot afford a 1.5-second pause has a different answer than a chat interface where users do not notice the difference. Faster inference does not automatically produce a business return; it produces faster inference, and its value is application-specific.

Sensitivity

  1. Model choice within a provider. Groq's Qwen 3.6 27B has a 5x output-to-input price ratio, so an output-heavy workload on that model can cost more than Llama 3.3 70B despite the smaller model size.
  2. Free-tier limits. Groq's free tier is rate-limited (30 requests/minute) in a way that blocks production traffic; Cerebras's free tier (1 million tokens/day) is more generous for batch-style testing.
  3. Throughput measurement method. Vendor-reported and independently benchmarked throughput for the same model can differ by 2x or more (Groq's Llama 3.3 70B: 293.6 tokens/second independently versus 403–750 on a vendor comparison page), which changes any latency-based value calculation substantially.
  4. Model substitution for cost control. Moving from Llama 3.3 70B to GPT-OSS 120B cuts Groq's cost per billion tokens by 64% while changing model capability, not just price.

Budgeting traps

  • Assuming faster automatically means better value. Cerebras costs more on every model in this set; its case rests entirely on whether the latency reduction is worth the premium for a specific application.
  • Pricing off the input rate alone. Output-heavy models (Qwen 3.6 27B on Groq, DeepSeek-R1 on SambaNova) can cost far more than the input rate suggests.
  • Free-tier rate limits mistaken for production capacity. Groq's 30 requests/minute free tier is a development sandbox, not a production allowance.
  • Comparing vendor-reported and independently benchmarked speed as if they were the same measurement.

What to ask before you buy

Ask each vendor for the exact per-token rate on the specific model and quantization you plan to run, and ask for their throughput under sustained production load rather than a best-case benchmark. Then decide whether your workload's latency sensitivity justifies Cerebras's premium over Groq's lower per-token price.


ARTICLE 6