NotCheapMAKE EVERY CREDIT COUNT

Vercel vs Cloudflare Workers vs AWS Lambda: Real Serverless Backend Cost for 100 Million AI API Requests in 2026

The short answer: these three platforms meter execution time in three genuinely different ways, and which one is cheapest flips entirely depending on how CPU-heavy the workload is. For a lightweight API proxy (short execution, low memory), AWS Lambda is the clear winner at about $19.80/month for 100 million requests, because its generous 400,000 GB-second free compute tier absorbs the entire workload's compute cost, leaving only the per-request charge. Cloudflare Workers comes in at $61.40 for the same shape, and Vercel at roughly $112, both of which bill CPU time or memory-time in ways that don't benefit from a free-tier threshold as generous relative to this workload. The moment the workload shifts to a heavier AI-orchestration shape — longer duration, more memory, genuinely CPU-bound — the ranking compresses dramatically: Lambda rises to about $846, Cloudflare Workers to $831, and Vercel to roughly $1,625, because a CPU-bound workload has no idle time for any platform's time-based billing to discount.

The billing-unit mismatch, up front

AWS Lambda bills wall-clock duration — GB-seconds, the product of allocated memory and total execution time, including any time spent idle waiting on a downstream call. Cloudflare Workers bills CPU time specifically — only the milliseconds the CPU is actually executing code, not time spent waiting on I/O, which can make Workers dramatically cheaper for I/O-bound workloads (a function waiting on a database or API response) and offers no such advantage for CPU-bound ones. Vercel's Fluid Compute splits the difference, billing Active CPU time (like Workers, compute-only) plus Provisioned Memory (GB-hours the function stays warm, closer to Lambda's wall-clock model) plus a per-seat platform fee layered on top of both. A workload that's mostly waiting — a typical lightweight API proxy calling a database — is billed very differently across these three; a workload that's genuinely CPU-bound the whole time it runs is billed much more similarly.

What each vendor bills

AWS Lambda (OFFICIAL, AWS Lambda pricing). $0.20 per million requests. $0.0000166667 per GB-second of allocated-memory-times-duration, including idle wait time within the invocation. Free tier: 1 million requests and 400,000 GB-seconds a month, permanently — not a 12-month trial, a standing allowance. At a common 256 MB/200 ms invocation shape, that free compute tier alone covers roughly 7.8 million invocations' worth of compute, well above the request-based free tier.

Cloudflare Workers (OFFICIAL, developers.cloudflare.com/workers pricing). Workers Paid: $5/month base fee, 10 million requests included, then $0.30 per additional million. 30 million CPU-milliseconds included per month, then $0.02 per additional million CPU-ms — billing only active compute time, with no charge for time spent waiting on network I/O, database calls, or downstream API responses. Free plan: 100,000 requests/day, no separate paid CPU allowance structure. Workers' architecture (V8 isolates, not full containers) also caps memory at 128 MB per isolate, a structural ceiling the other two don't share at this tier.

Vercel (OFFICIAL, current Vercel pricing documentation). Pro: $20 per seat per month, which includes a matching $20 of usage credit per seat — the credit offsets metered usage dollar-for-dollar up to that amount, after which usage bills separately on top of the seat fee. Function Invocations: $0.60 per million, with 1 million included. Active CPU: $0.128 per CPU-hour (one source reports this varies $0.128–$0.221 by region), with 4 CPU-hours included. Provisioned Memory (GB-hours while a function stays warm): $0.0106 per GB-hour, with 360 GB-hours included. A separate, materially higher $0.18-per-GB-hour figure appears in some older guides; that figure belongs to a prior, superseded pricing context and does not reflect Vercel's current rate. Vercel requires at least one paid seat to access Pro-tier metered usage at all, meaning the $20 seat fee is a genuine floor, not just a credit mechanism.

Two workload shapes

All ILLUSTRATIVE. Shape A: lightweight API proxy — 128 MB memory, 20 ms total duration, roughly 15 ms of actual CPU time (the rest idle, waiting on a fast downstream call). Shape B: heavier AI orchestration request — 1 GB memory, 500 ms total duration, roughly 400 ms of actual CPU time (closer to continuously CPU-bound, representative of in-process model-adjacent work like parsing, validation, or light inference orchestration).

Cost at 100 million requests/month, by workload shape

Formula: Lambda = (requests ÷ 1M × $0.20, net of 1M free) + (GB-seconds × $0.0000166667, net of 400,000 free); Workers = $5 base + (requests ÷ 1M × $0.30, net of 10M free) + (CPU-ms ÷ 1M × $0.02, net of 30M free); Vercel = $20 seat + max(0, metered usage − $20 credit), where metered usage sums invocations, Active CPU, and Provisioned Memory charges net of their respective free allowances.

PlatformShape A: lightweight proxyShape B: heavy AI orchestration
AWS Lambda$19.80 (compute fully covered by free tier)$846.47 ($19.80 requests + $826.67 compute)
Cloudflare Workers$61.40 ($5 base + $27.00 requests + $29.40 CPU)$831.40 ($5 base + $27.00 requests + $799.40 CPU)
Vercel$112.22 ($59.40 invocations + $52.82 CPU + $0 memory, within free allowance)$1,624.52 ($59.40 invocations + $1,421.71 CPU + $143.41 memory)

Lambda's advantage at the lightweight shape comes entirely from its free-tier generosity relative to a low-memory, short-duration workload — not from a structurally cheaper rate. At the heavy shape, Lambda and Workers converge to within 2% of each other, because both platforms' free allowances become comparatively small relative to a genuinely CPU-bound 100-million-request workload, and the underlying per-unit rates end up mattering far more than which platform's free tier is more generous.

Why idle time is the real variable, not request count

Formula: idle-time cost advantage = wall-clock billed time − CPU-only billed time, which is zero on Lambda and Vercel's memory component (both bill wall-clock-adjacent time) but can be substantial on Workers and Vercel's Active CPU component. A function spending 180 ms waiting on a slow downstream API call, with only 8 ms of actual CPU work, pays for all 180 ms on Lambda (via GB-seconds) but only the 8 ms on Workers — a gap that widens the more I/O-bound the workload is, and vanishes entirely for a function that's CPU-bound the whole time it runs, which is exactly why Shape B's results converge so much more tightly than Shape A's.

Sensitivity

  1. CPU-bound versus I/O-bound workload shape. The single largest lever in this entire comparison; it is what separates the wide spread at Shape A from the narrow spread at Shape B.
  2. Memory-heavy versus compute-light function design on Vercel. Provisioned Memory and Active CPU are billed separately; a function that stays warm with a large memory allocation but modest CPU use shifts cost toward the memory line, and vice versa.
  3. Memory allocation. Directly and linearly scales Lambda's GB-second cost and Vercel's Provisioned Memory cost; Workers is unaffected by memory allocation since it bills CPU time only, within its fixed 128 MB isolate ceiling.
  4. Seat count on Vercel. Each additional seat adds another $20 base fee and another $20 credit; a team provisioning more seats than needed for billing purposes alone pays for credits it may not fully use.

Budgeting traps

  • Assuming a lightweight workload's relative platform ranking holds for a heavier one. Lambda's advantage at Shape A almost entirely evaporates at Shape B, where all three platforms land within a comparable range.
  • Citing an older $0.18-per-GB-hour Provisioned Memory figure for Vercel. That rate belongs to a superseded pricing context; the current confirmed rate is $0.0106 per GB-hour.
  • Treating Cloudflare Workers' 128 MB per-isolate memory ceiling as a non-issue. A workload genuinely requiring more memory per invocation cannot run on Workers' standard isolate model at all, regardless of price.
  • Forgetting that Lambda's free tier is measured in GB-seconds, not a flat invocation count. A low-memory, short-duration workload can process far more invocations within the free allowance than a naive "1 million free requests" reading would suggest.

What to ask before you buy

Measure your actual workload's CPU-bound versus idle-time ratio before comparing these three platforms, since that ratio — not raw request count — is what determines which platform's billing model is actually cheaper. Confirm Vercel's current rates directly against its live pricing page before committing to a multi-month budget, since serverless platform pricing in this category has moved more than once in recent history.