Baseten vs Modal vs RunPod Serverless: Real GPU Model Serving Cost per 1 Million Inference Runs in 2026
The short answer: on the bare hourly GPU rate, Baseten is only modestly more expensive than Modal or RunPod, not dramatically so. On an identical H100, Modal bills $3.95/hour and RunPod Serverless bills roughly $3.49/hour, both metered per second of actual compute. Baseten lists $6.50/hour, about 65–86% higher — but Baseten prices Dedicated Deployment replica compute down to the minute, where autoscaling controls how many replicas stay running, rather than billing each individual inference request as its own rounded-up unit. As a theoretical full-utilization compute floor — the total GPU-hours the work itself requires, with no idle replica time — 1 million calls averaging 2 seconds each cost roughly $1,939 on RunPod, $2,194 on Modal, and $3,611 on Baseten: Baseten runs about 1.6 to 1.9 times the other two on the same underlying compute, not an order of magnitude more. What a real Baseten invoice actually comes to depends heavily on autoscaling behavior, concurrency, and how long idle replicas stay warm before scaling down — none of which this article can derive into one normalized per-request rate without first fixing those assumptions.
What each vendor bills
Modal (OFFICIAL, Modal pricing documentation). Per-second billing, no charge for idle time between calls: H100 SXM: $3.95/hour ($0.001097/second). A100 80GB: $2.50/hour ($0.000694/second). CPU billed separately at $0.0000131/core-second, memory at $0.00000222/GiB-second. No minimum billing increment beyond the second.
RunPod Serverless (REPORTED, cross-referenced across multiple 2026 comparison sources). Also per-second billing on its serverless product (distinct from RunPod's on-demand dedicated-pod pricing, which bills hourly): H100: approximately $3.49/hour ($0.0009694/second). A100: approximately $1.59/hour ($0.0004417/second), both reported as RunPod's serverless-specific rates rather than its separately listed on-demand instance pricing ($2.39–$2.69/hour for H100 on-demand, which is a different product).
Baseten (OFFICIAL for the billing mechanism; REPORTED for the exact rate). H100: approximately $6.50/hour. Baseten's own first-party material describes pricing its Dedicated Deployment replica compute down to the minute, not a separate 60-second charge applied independently to every individual request. In practice, an inference endpoint on Baseten runs as one or more persistent replicas that autoscaling spins up and down based on traffic; the minute-level granularity applies to how long each replica stays running and billable, not to each 2-second request in isolation. A free tier reportedly provides 10 hours of GPU time.
The aggregate compute floor, as a theoretical baseline
Formula: total GPU-hours = runs × GPU-seconds per run ÷ 3,600. This is the minimum possible compute cost on any of these three platforms — it assumes perfect, continuous utilization with zero idle replica time, which no real production workload achieves exactly, but it is the right baseline for comparing the three vendors' underlying hourly rates on equal footing before autoscaling overhead is added.
| Call length | Total GPU-hours (1M runs) | Modal ($3.95/hr) | RunPod (~$3.49/hr) | Baseten ($6.50/hr) |
|---|---|---|---|---|
| 2 seconds | 555.6 | $2,194 | $1,939 | $3,611 |
| 10 seconds | 2,777.8 | $10,972 | $9,694 | $18,056 |
| 60 seconds | 16,666.7 | $65,833 | $58,167 | $108,333 |
At this theoretical floor, Baseten runs roughly 1.6 to 1.9 times Modal's or RunPod's cost at every call length tested — a real, meaningful premium, but nowhere near an order of magnitude, and nothing like treating each short call as an independent 60-second billing event would suggest.
What this floor leaves out, and why a real invoice needs more assumptions
The table above assumes zero idle replica time, which is not how autoscaled serverless deployment actually works on any platform. A real Baseten (or Modal, or RunPod) bill depends on several factors this article cannot fix without inventing workload-specific assumptions:
- Autoscaling replica count. How many replicas an endpoint runs at a given moment, driven by concurrent traffic, not by request count alone.
- Concurrency and throughput per replica. A single replica serving multiple concurrent requests spreads its billed time across more inference work than one handling requests sequentially.
- Scale-to-zero behavior and the autoscaling window. How quickly a platform scales replicas down after traffic drops, and whether it scales to zero at all between bursts.
- Idle replica minutes. Time a replica stays warm and billable between requests, waiting for the next one, which is not "zero" in any real deployment and is where the gap between the theoretical floor above and an actual invoice opens up.
A fully normalized, confident cost-per-1-million-runs figure that accounts for all four of these cannot be derived from public information without first fixing a specific concurrency and traffic-shape assumption for each platform's autoscaling behavior — doing so here would mean inventing a per-request rounding model that does not reflect how Baseten's first-party material describes its own billing, which is exactly the error this repair corrects. What can be said with confidence is the floor above (roughly 1.6–1.9x) and the direction of the real-world gap: a bursty, low-concurrency workload with long idle gaps between requests will push all three platforms' real cost above this floor, and is likely to widen Baseten's premium over Modal and RunPod rather than narrow it, since idle replica minutes bill at Baseten's higher hourly rate the same as active ones.
Why Baseten's model makes sense for a different workload shape
Baseten explicitly positions itself around managed, production-grade model serving rather than raw general-purpose serverless compute — its Dedicated Deployment model reflects a workload where models run as persistent, higher-utilization endpoints rather than invoked as short, sparse function calls. For a workload with sustained, high-concurrency traffic that keeps replicas consistently busy, the gap between Baseten and Modal/RunPod should stay close to the theoretical floor's 1.6–1.9x, and Baseten's managed-serving feature set (autoscaling policy controls, dashboards, alerting, a Truss-based packaging framework) may offset that remaining premium for teams that value not managing that infrastructure themselves. For a bursty, low-traffic, low-concurrency workload, idle replica time is likely to push Baseten's real cost further above Modal's and RunPod's per-second-metered floor.
Sensitivity
- Concurrency and replica utilization. The dominant unmodeled variable in this article; a high-concurrency workload that keeps replicas consistently busy stays close to the theoretical floor, while a low-concurrency, bursty one pushes real cost meaningfully higher, especially on Baseten.
- Which GPU tier. The relative ranking (RunPod cheapest, Modal second, Baseten priciest per-hour) holds across the A100 and H100 tiers checked, though the exact dollar gap changes with GPU class.
- Autoscaling scale-down speed. A platform or configuration that scales replicas down quickly after traffic drops minimizes idle-time billing; a slow scale-down window extends it.
- Cold-start overhead. Not modeled in this article's cost figures, but relevant alongside autoscaling behavior — a platform with a slower cold start may need to keep more replicas warm (and billable) to maintain latency, raising real cost beyond the theoretical floor.
Budgeting traps
- Treating Baseten's per-minute billing as a flat per-request rounding charge. Baseten's own material describes minute-level billing on replica/deployment compute time, not a separate 60-second charge applied independently to every short request.
- Assuming the theoretical full-utilization floor in this article is what you'll actually be billed. Real cost depends on concurrency and autoscaling behavior this article cannot fix without inventing workload-specific assumptions.
- Assuming RunPod's on-demand dedicated-instance rate applies to its Serverless product. The two are different pricing structures on the same platform.
- Sizing a low-concurrency, bursty workload onto Baseten without accounting for idle replica time. The managed feature set may still be worth it, but the real cost gap versus Modal or RunPod is likely wider than the bare hourly-rate comparison suggests.
What to ask before you buy
Ask Baseten directly for guidance on expected replica count and idle-time behavior for your specific traffic pattern and concurrency level, since a confident per-request cost cannot be derived from the published hourly rate alone. Confirm RunPod's current Serverless-specific GPU rate directly, since it is a different, generally cheaper product than RunPod's on-demand dedicated-instance pricing that shares the same brand name.