Modal vs RunPod Serverless vs Replicate: Real AI Inference Hosting Cost in 2026
Rates are USD list prices. Modal figures are GPU-only unless a row says otherwise: Modal bills CPU and memory separately, while the RunPod pod rates include bundled vCPUs and RAM. RunPod's pricing page is dated 13 September 2026; a comparison of Secure Cloud pods is used (Community Cloud is cheaper and less reliable). Modal and Replicate rates as published in September 2026.
The short answer
At low volume, per-second platforms cost tens of dollars while an always-on A100 costs $1,160.70 a month at the current $1.59 rate. For steady traffic on one A100 at 60% utilization, Modal's GPU-only charge is $1,094.30 before overhead (5.7% below the pod) and $1,203.73 with 10% start and idle overhead (3.7% above it); RunPod Serverless costs $1,191.36 and $1,310.50; a Replicate private deployment kept online costs $3,679.20. For two H100s at 80% utilization, rented PCIe pods win ($4,219.40) over Modal ($4,612.67) and RunPod Serverless ($5,594.72). No platform is cheapest everywhere; the pod-versus-serverless line sits at 58% to 73% utilization.
Rates per active hour
| GPU | Modal (per second) | RunPod Serverless (flex) | RunPod pod (Secure Cloud) | Replicate |
|---|---|---|---|---|
| A100 80GB | $0.000694/s = $2.498/hr | $2.72/hr | $1.59/hr (PCIe and SXM) | $0.0014/s = $5.04/hr |
| H100 | $0.001097/s = $3.949/hr | $4.79/hr (was $4.55) | $2.89 PCIe, $3.49 SXM | $5.49/hr |
| L4 / L40S | $0.80 / $1.95 per hr | Not itemized | Not itemized | L40S $3.51/hr |
RunPod's A100 PCIe pod was $1.39 until its September 2026 refresh. Modal adds CPU ($0.0000131 per core-second) and memory ($0.00000222 per GiB-second). RunPod Serverless bills from worker start to stop, including start time and an idle timeout that defaults to 5 seconds; network storage starts at $0.05/GB/month. Replicate public models bill only active time; private models and deployments bill setup, idle and active time.
Modal with CPU and memory (illustrative configurations)
Modal's CPU is $0.0000131 per core-second and memory $0.00000222 per GiB-second (reported). For an illustrative A100 container of 4 cores and 32 GiB, the extra is 4 × 0.0000131 + 32 × 0.00000222 = $0.00012344 per second, or $0.444 per hour; for an H100 container of 8 cores and 64 GiB it is $0.00024688 per second, or $0.889 per hour. With these, Modal costs $2.94 per A100 hour and $4.84 per H100 hour, and the break-even utilizations against the pods become 54.0% (A100, $1.59 ÷ $2.94) and 59.7% (H100 PCIe, $2.89 ÷ $4.84). The RunPod pod hours below include their bundled compute; other CPU and memory choices change these figures.
Workload A: bursty, low volume
Billed time is defined differently by billing model, so the rows below are not equal compute. Assumptions: 2,000 requests a month on an A100 80GB; 5 seconds of execution, 4 seconds of amortized start and model load, 5 seconds of idle tail per request, so 14 billed seconds per request or 28,000 seconds (7.78 hours). Start time is an assumption, not a measured figure.
| Platform | Formula | Monthly |
|---|---|---|
| Modal | 28,000 × $0.000694 | $19.43 |
| RunPod Serverless | 7.78 hr × $2.72 | $21.16 |
| Replicate, public model | 10,000 s × $0.0014 (setup and idle free) | $14.00 |
| Replicate, private model | 28,000 s × $0.0014 | $39.20 |
| RunPod pod, always on | 730 hr × $1.59 | $1,160.70 |
| Modal with 4 cores + 32 GiB | 28,000 s × ($0.000694 + $0.00012344) | $22.89 |
Equal compute, execution only: if each platform billed only the 10,000 seconds of execution, Modal (GPU-only) would cost 10,000 × $0.000694 = $6.94, RunPod Serverless 2.78 hr × $2.72 = $7.56 and Replicate $14.00. Replicate's public-model row bills only that active time; Modal and RunPod bill container or worker time including start and idle, and Replicate private models bill setup, idle and active time (28,000 seconds, $39.20). The $14.00 public-model figure is therefore not directly comparable with the 28,000-second rows.
Workload B: steady, one A100 at 60% utilization
438 active hours in a 730-hour month.
| Platform | Zero overhead | With 10% overhead |
|---|---|---|
| Modal, GPU-only | 438 × $2.4984 = $1,094.30 | $1,203.73 |
| Modal with 4 cores + 32 GiB | 438 × $2.9428 = $1,288.94 | $1,417.83 |
| RunPod pod | 730 × $1.59 = $1,160.70 | $1,160.70 |
| RunPod Serverless | 438 × $2.72 = $1,191.36 | $1,310.50 |
| Replicate private, one instance online | 730 × $5.04 = $3,679.20 | $3,679.20 |
Break-even utilization = pod rate ÷ serverless rate: $1.59 ÷ $2.4984 = 63.6% (Modal, GPU-only) or 54.0% with the CPU and memory above; $1.59 ÷ $2.72 = 58.5% (RunPod Serverless). At 60% the pod is cheaper than RunPod Serverless but not than Modal until overhead is counted.
Workload C: two H100s at 80% utilization
1,168 active GPU-hours; 1,460 pod-hours.
| Platform | Formula | Monthly |
|---|---|---|
| RunPod pods (PCIe) | 1,460 × $2.89 | $4,219.40 |
| Modal, GPU-only | 1,168 × $3.9492 | $4,612.67 ($4,751.05 with 3% retry loss) |
| Modal with 8 cores + 64 GiB | 1,168 × $4.8380 | $5,650.75 |
| RunPod pods (SXM) | 1,460 × $3.49 | $5,095.40 |
| RunPod Serverless | 1,168 × $4.79 | $5,594.72 |
| Replicate, active time | 1,168 × $5.49 | $6,412.32 |
| Replicate, two deployments online | 1,460 × $5.49 | $8,015.40 |
H100 break-even utilization: Modal (GPU-only) versus a PCIe pod $2.89 ÷ $3.9492 = 73.2% (59.7% with CPU and memory); versus an SXM pod $3.49 ÷ $3.9492 = 88.4%; RunPod Serverless versus a PCIe pod $2.89 ÷ $4.79 = 60.3%. PCIe and SXM cards differ in interconnect, so the pod row is not a like-for-like comparison with every serverless H100.
What the advertised price misses
Idle hours are the main cost of a pod. Cold starts bill at the full per-second rate on RunPod Serverless, and model loading time depends on model size and platform configuration; no measured start time is assumed as a vendor fact. RunPod lists no data ingress or egress charge. Replicate's public-versus-private difference means a workload priced at $14 a month on a public model can cost $39 privately at low volume and $3,679 if kept warm.
Which workload fits
Spiky traffic and prototypes: Modal or RunPod Serverless, or Replicate public models where a suitable public model exists. Steady traffic near 60%: compare Modal and a pod using measured overhead. Sustained loads above roughly 60–75%: rented pods, sized to load.
Sources
Modal: modal.com/pricing, rates as read by Markaicode, 3 Sep 2026. RunPod: runpod.io/pricing (updated 13 Sep 2026), serverless billing docs; rates as verified by UsagePricing, 22 Sep 2026 and Omid Saffari, 25 Sep 2026. Replicate: replicate.com/pricing; hardware rates as reported by Markaicode, Aug 2026 and CheckThat.