NotCheapMAKE EVERY CREDIT COUNT

Cloudflare Workers AI vs Groq vs Together AI: Real Serverless LLM Cost per 1 Billion Tokens in 2026

The short answer: Groq is the cheapest of the three at the 8B parameter tier, undercutting Cloudflare Workers AI's cheapest fp8-optimized model variant by a wide margin and Together AI by even more. At 1 billion tokens (5:1 input:output ratio) on an 8B-class model, Groq costs about $55, against $102 on Cloudflare's fp8-fast variant and $180 on Together AI. At the 70B tier, the picture changes: Cloudflare ($620) and Groq ($623) are within 0.5% of each other, and Together AI ($880) remains the most expensive of the three at both tiers tested. Cloudflare's competitiveness depends heavily on calling its cheapest quantized model variant, which most buyers miss. Cloudflare is also mid-transition on its billing mechanism, moving away from its opaque "Neurons" compute unit toward simpler per-token pricing tiered by model size — both systems appear in its current documentation.

What each vendor bills

Cloudflare Workers AI (OFFICIAL, developers.cloudflare.com/workers-ai/platform/pricing). Every model's price converts internally to Neurons, an AI-specific compute unit, billed at $0.011 per 1,000 Neurons, with 10,000 free Neurons per day resetting daily. The Neuron cost varies sharply by model — the published per-token dollar rate is derived from that Neuron figure, not set independently. Example current rates: llama-3.1-8b-instruct-fp8-fast: $0.045 input / $0.384 output per million tokens. llama-3.1-8b-instruct (unquantized): $0.282 / $0.827 — over 6x the fp8-fast variant's input price for what is nominally "the same" model, because quantization changes the Neuron cost per token substantially. llama-3.3-70b-instruct-fp8-fast: $0.293 / $2.253. Cloudflare has separately announced it is moving away from Neurons toward a simpler model-size-tiered pricing scheme (models ≤3B: $0.10/M blended tokens; 3.1B–8B: $0.15/M; 8.1B–20B: $0.20/M; 20.1B–40B: $0.50/M), described in a company blog post using future tense ("we will be moving towards"); as of the documentation reviewed, the per-model Neuron-derived rate table above is still the one actually published and billing, so this article uses it as current, while flagging the announced transition as a near-term change worth confirming before locking in a long-term budget.

Groq (OFFICIAL, Groq pricing page, as established in this site's companion high-speed-inference comparison). Llama 3.1 8B Instant: $0.05 input / $0.08 output per million tokens. Llama 3.3 70B Versatile: $0.59 / $0.79. Groq's rates are flat regardless of quantization variant naming, since Groq runs its own curated model set on custom LPU silicon rather than offering multiple quantization levels of the same open-weight model as separate SKUs the way Cloudflare does.

Together AI (OFFICIAL, Together AI pricing/model pages, as established in this site's companion inference-cost comparison). Llama 3.1 8B: $0.18 input / $0.18 output per million tokens (flat). Llama 3.3 70B: $0.88 / $0.88 (flat). Together prices input and output identically on these specific models, unlike Cloudflare's and Groq's output-weighted rates.

Cost per 1 billion tokens, by model tier

Formula: cost per 1B = (5 × input rate + output rate) ÷ 6 × 1,000, at an ILLUSTRATIVE 5:1 input:output ratio.

TierCloudflare (fp8-fast variant)GroqTogether AI
8B-class$102$55$180
70B-class$620$623$880

At the 8B tier, Groq is the cheapest by a wide margin (46% below Cloudflare, 69% below Together), driven mostly by Groq's very low output-token rate ($0.08/M, less than a quarter of Cloudflare's fp8-fast output rate). At the 70B tier, that gap closes almost entirely: Cloudflare and Groq are essentially tied, and Together AI's flat, un-discounted-for-output pricing structure leaves it the most expensive of the three at both tiers.

Volume scenarios

Formula: monthly cost = input millions × input rate + output millions × output rate, at 100 million input / 20 million output and 1 billion input / 200 million output.

TierVolumeCloudflareGroqTogether AI
8B-class100M in / 20M out$12.18/mo$6.60/mo$21.60/mo
8B-class1B in / 200M out$121.80/mo$66.00/mo$216.00/mo
70B-class100M in / 20M out$74.36/mo$74.80/mo$105.60/mo
70B-class1B in / 200M out$743.60/mo$748.00/mo$1,056.00/mo

The quantization-variant trap on Cloudflare

This is the single most consequential thing to get right when budgeting Cloudflare Workers AI, and it does not exist as a comparable issue on Groq or Together. On the input-token rate alone, Cloudflare's unquantized llama-3.1-8b-instruct costs $0.282, about 6.3 times its own llama-3.1-8b-instruct-fp8-fast variant at $0.045 — nominally the "same" model, at wildly different input prices depending on which exact model string is called. That 6.3x figure is input-only, though: under this article's modeled 5:1 input:output ratio, the blended cost per billion tokens is $373 unquantized versus $102 for fp8-fast, an overall difference of about 3.7x (output tokens are relatively closer between the two variants than input tokens are, which narrows the blended gap below the input-only ratio). A team that copies a model name from an old code sample or documentation page without checking for a newer, cheaper quantized variant meaningfully overpays on Cloudflare for identical model capability — roughly 3.7x on a typical blended workload, and more on an input-heavy one — a trap that does not exist in the same form on Groq (single curated rate per model) or Together (one rate per named model, quantization handled server-side without separate SKUs).

Sensitivity

  1. Which exact Cloudflare model variant is called. As shown above, this alone is roughly a 3.7x swing at the 8B tier on a blended 5:1 workload, and up to 6.3x on an input-heavy one.
  2. Output-heavy versus input-heavy workloads. Groq's and Cloudflare's output rates are 1.6–8.5x their input rates depending on model; a summarization-style, output-light workload favors these platforms more than the 5:1 ratio used here suggests, and a generation-heavy workload favors them less.
  3. Model tier. The relative ranking of these three platforms is not stable across model sizes — Groq's large 8B-tier advantage nearly disappears at the 70B tier.
  4. Cloudflare's pending pricing-model transition. The announced move away from Neurons toward flat size-tiered pricing would materially change Cloudflare's numbers in this article once it takes effect; reconfirm the live rate before budgeting a multi-month commitment.

Budgeting traps

  • Calling an old or default Cloudflare model string instead of its cheapest current quantized variant. This is the largest, most avoidable cost error in this comparison.
  • Assuming Groq's large 8B-tier price advantage holds at every model size. It does not; the 70B tier is close to a three-way tie.
  • Ignoring Cloudflare's free daily Neuron allocation when estimating a small-scale bill. 10,000 Neurons a day, reset daily, covers meaningful volume on cheaper model variants before any billing starts.
  • Budgeting Cloudflare off its announced future pricing scheme rather than its currently live, published rate table. The two coexist in current documentation; only the Neuron-derived per-model table is confirmed to be billing today.

What to ask before you buy

Confirm you are calling the cheapest available quantized variant of your target model on Cloudflare specifically, since the same nominal model can be billed at wildly different rates depending on the exact model string. Ask Cloudflare directly whether its announced size-tiered pricing has gone live before locking in a long-term budget based on today's per-model rate table.