Portkey vs Helicone vs Cloudflare AI Gateway: Real LLM Gateway Cost per 1 Million Requests in 2026
The short answer: a gateway's invoice is usually the smallest number in this comparison, and the gap between a user request and an upstream model request is the number that actually matters. A router with a retry, a fallback and a cache in front of it can turn one user request into 1.0 to 2.04 upstream model calls depending on shape, and every one of those calls is billed by the model provider regardless of what the gateway charges. On raw gateway fees, Portkey's Production plan costs $49 to $130 per 1 million logged requests depending on volume, and its logging fee alone breaks even against a mid-tier model's per-call cost once caching or routing avoids as little as 2.3% to 24.5% of upstream calls, depending on request shape and volume. Both Portkey and Helicone changed ownership in 2026, which matters for anyone evaluating their roadmaps, not just their price sheets.
Vendor status and what each one bills
Portkey (OFFICIAL for pricing, REPORTED for the acquisition). Dev is free with 10,000 logs a month; Production is $49 a month with 100,000 logs included, then $9 per additional 100,000 logs; Enterprise is custom, unlimited requests, 90-day retention (QUOTE-ONLY / UNKNOWN). Palo Alto Networks announced its acquisition of Portkey on April 30, 2026 and closed the deal on May 29, 2026; Portkey now operates inside Palo Alto's Prisma AIRS platform (REPORTED). The pricing page does not publish an overage rate beyond the $9-per-100,000-logs block, so budgeting above the Enterprise threshold requires a quote.
Helicone (OFFICIAL for pricing, REPORTED for the acquisition). Hobby is free with 10,000 requests a month and 7-day retention; Pro is $79 a month with 10,000 requests free, then usage-based; Team is $799 a month with 10,000,000 requests free, then usage-based; Enterprise is custom with configurable or forever retention. Mintlify acquired Helicone on March 3, 2026, and the product is now in maintenance mode: security updates and model support continue, but active feature development has stopped (REPORTED). Helicone's per-request overage rate beyond the free allotments on Pro and Team is not published (QUOTE-ONLY / UNKNOWN); its own pricing calculator returns estimated totals without showing the underlying per-unit rate. Helicone's separate AI Gateway product, launched September 2025, passes through model costs at a 0% markup, so gateway routing itself adds no model-side surcharge on that product line.
Cloudflare AI Gateway (OFFICIAL). Free, bundled with a Cloudflare Workers plan, with logging, caching, rate limiting and fallback built in. There is no separate per-request AI Gateway charge; cost instead follows whatever Cloudflare Workers plan the account already carries. This makes Cloudflare structurally different from the other two: its gateway fee is effectively $0 on top of infrastructure most Cloudflare-using teams already pay for, but it is not a standalone product with its own price ladder to compare unit-for-unit.
Normalizing the meter: a user request is not a model request
Three shapes, all ILLUSTRATIVE, describe how many upstream model calls one user-facing request generates:
| Shape | Description | Upstream model calls per user request |
|---|---|---|
| A: direct single-model request | No retry, no fallback | 1.0 |
| B: one retry or fallback | A failed or slow primary call triggers one fallback call on 50% of requests | 1.5 |
| C: router with cache, fallback and guardrail | A cache miss (80% of requests, 0.80 calls), a fallback on 30% of those misses (0.24 calls), and a separate guardrail-check call on every request (1.00 call) | 2.04 |
Formula: upstream calls = user requests × multiplier, where the multiplier is the sum of its components (0.80 + 0.24 + 1.00 = 2.04 for shape C). At 1 million user requests a month, shape C generates 2.04 million upstream model calls, all billed by the model provider, while the gateway itself is typically billed on the user request or logged request count, not the upstream count.
Gateway cost per 1 million user requests (Portkey)
Formula: Portkey cost = $49 + ceiling(max(0, logs − 100,000) ÷ 100,000) × $9, applied to logged (user-facing) requests, flat across shapes A, B and C because Portkey's meter is logs, not upstream calls.
| User requests/month | Portkey cost | Per 1 million user requests |
|---|---|---|
| 100,000 | $49 | $490 |
| 1,000,000 | $130 | $130 |
| 10,000,000 | $940 | $94 |
| 100,000,000 | $9,040 | $90.40 |
The per-million rate falls sharply between 100,000 and 1 million requests because the $49 base fee is amortized over a larger volume, then flattens as the linear $9-per-100,000-logs overage takes over.
Helicone and Cloudflare: what can and cannot be priced
Helicone's Team plan covers 10 million requests a month for a flat $799, or $79.90 per 1 million requests at exactly that volume, with no published rate for exceeding it. Below 10 million requests, Helicone's usage-based overage past the Pro plan's 10,000 free requests has no public per-request rate, so a precise figure at 100,000 or 1 million requests is QUOTE-ONLY / UNKNOWN; only the $79 (Pro) or $799 (Team) plan floor is known. Cloudflare's AI Gateway has no separate meter to convert into a per-million-request figure: its cost is whatever Workers plan and request volume the account already carries for reasons unrelated to the AI Gateway feature.
Total upstream model requests generated, and the amplification factor
| Shape | User requests/month | Upstream model calls generated | Amplification |
|---|---|---|---|
| A | 1,000,000 | 1,000,000 | 1.0x |
| B | 1,000,000 | 1,500,000 | 1.5x |
| C | 1,000,000 | 2,040,000 | 2.04x |
A gateway that adds a retry or a fallback without also adding caching can materially increase the model bill even while its own logging fee stays flat, because the gateway meters logged requests while the model provider meters every call that actually reaches it.
Savings required from caching and routing to break even
Formula: break-even cache/routing hit rate = gateway fee ÷ (upstream calls × cost per model call), using an ILLUSTRATIVE mid-tier model cost of $0.002 per call (a blended estimate across input and output tokens for a short request).
| Shape | Volume | Gateway fee | Upstream calls | Break-even hit rate |
|---|---|---|---|---|
| A | 100,000 | $49 | 100,000 | 24.5% |
| A | 1,000,000 | $130 | 1,000,000 | 6.5% |
| A | 10,000,000 | $940 | 10,000,000 | 4.7% |
| B | 100,000 | $49 | 150,000 | 16.3% |
| B | 1,000,000 | $130 | 1,500,000 | 4.3% |
| B | 10,000,000 | $940 | 15,000,000 | 3.1% |
| C | 100,000 | $49 | 204,000 | 12.0% |
| C | 1,000,000 | $130 | 2,040,000 | 3.2% |
| C | 10,000,000 | $940 | 20,400,000 | 2.3% |
At any meaningful volume, a gateway's own fee is cheap to recover through caching: even a modest 2–5% cache hit rate at 1 million or more requests a month covers the entire Portkey Production bill. The real cost driver in this category is not the gateway invoice; it is whether the routing logic actually reduces upstream calls or, through retries and fallback chains, increases them.
Sensitivity
- Retry and fallback rate. Raising shape B's fallback trigger rate from 50% to 80% of failed calls pushes the multiplier from 1.5 to about 1.8, proportionally increasing upstream model spend without changing the gateway fee.
- Model cost per call. A frontier model at 5–10 times the assumed $0.002 per call lowers the break-even hit rate proportionally; caching becomes even more valuable on expensive models.
- Plan floor versus volume. Portkey's per-million rate is dominated by the $49 floor below 1 million requests and by the linear $9-per-100,000-logs rate above it.
- Helicone's unknown overage. Any team above the Pro plan's small free allotment and below Team's 10-million-request ceiling cannot get a precise number from public sources; budgeting there requires contacting Helicone directly.
Budgeting traps
- Treating "requests" as one meter. A gateway's logged requests, the upstream calls it generates, and the underlying tokens billed by the model provider are three different numbers.
- Retries without visibility. A fallback chain that silently doubles upstream calls will not show up as a gateway cost increase; it shows up on the model provider's invoice.
- Post-acquisition uncertainty. Both Portkey (now under Palo Alto Networks) and Helicone (now under Mintlify, in maintenance mode) changed hands in 2026, which affects roadmap and support expectations independent of today's price sheet.
- Unpublished overage rates. Helicone's per-request rate beyond its free allotments is not public; do not assume it matches a competitor's rate.
What to ask before you buy
Ask each vendor for the price at your logged-request volume and, separately, for what share of your traffic today involves a retry, fallback or cache lookup. Multiply your user-request volume by that shape's amplification factor to estimate the real number of upstream model calls before pricing the model side of the bill.
ARTICLE 4