NotCheapMAKE EVERY CREDIT COUNT

Modal vs RunPod Serverless vs Replicate: Real AI Inference Hosting Cost in 2026

Rates are USD list prices. Modal figures are GPU-only unless a row says otherwise: Modal bills CPU and memory separately, while the RunPod pod rates include bundled vCPUs and RAM. RunPod's pricing page is dated 13 September 2026; a comparison of Secure Cloud pods is used (Community Cloud is cheaper and less reliable). Modal and Replicate rates as published in September 2026.

DIRECT ANSWER

The short answer

At low volume, per-second platforms cost tens of dollars while an always-on A100 costs $1,160.70 a month at the current $1.59 rate. For steady traffic on one A100 at 60% utilization, Modal's GPU-only charge is $1,094.30 before overhead (5.7% below the pod) and $1,203.73 with 10% start and idle overhead (3.7% above it); RunPod Serverless costs $1,191.36 and $1,310.50; a Replicate private deployment kept online costs $3,679.20. For two H100s at 80% utilization, rented PCIe pods win ($4,219.40) over Modal ($4,612.67) and RunPod Serverless ($5,594.72). No platform is cheapest everywhere; the pod-versus-serverless line sits at 58% to 73% utilization.

Rates per active hour

GPUModal (per second)RunPod Serverless (flex)RunPod pod (Secure Cloud)Replicate
A100 80GB$0.000694/s = $2.498/hr$2.72/hr$1.59/hr (PCIe and SXM)$0.0014/s = $5.04/hr
H100$0.001097/s = $3.949/hr$4.79/hr (was $4.55)$2.89 PCIe, $3.49 SXM$5.49/hr
L4 / L40S$0.80 / $1.95 per hrNot itemizedNot itemizedL40S $3.51/hr

RunPod's A100 PCIe pod was $1.39 until its September 2026 refresh. Modal adds CPU ($0.0000131 per core-second) and memory ($0.00000222 per GiB-second). RunPod Serverless bills from worker start to stop, including start time and an idle timeout that defaults to 5 seconds; network storage starts at $0.05/GB/month. Replicate public models bill only active time; private models and deployments bill setup, idle and active time.

Modal with CPU and memory (illustrative configurations)

Modal's CPU is $0.0000131 per core-second and memory $0.00000222 per GiB-second (reported). For an illustrative A100 container of 4 cores and 32 GiB, the extra is 4 × 0.0000131 + 32 × 0.00000222 = $0.00012344 per second, or $0.444 per hour; for an H100 container of 8 cores and 64 GiB it is $0.00024688 per second, or $0.889 per hour. With these, Modal costs $2.94 per A100 hour and $4.84 per H100 hour, and the break-even utilizations against the pods become 54.0% (A100, $1.59 ÷ $2.94) and 59.7% (H100 PCIe, $2.89 ÷ $4.84). The RunPod pod hours below include their bundled compute; other CPU and memory choices change these figures.

Workload A: bursty, low volume

Billed time is defined differently by billing model, so the rows below are not equal compute. Assumptions: 2,000 requests a month on an A100 80GB; 5 seconds of execution, 4 seconds of amortized start and model load, 5 seconds of idle tail per request, so 14 billed seconds per request or 28,000 seconds (7.78 hours). Start time is an assumption, not a measured figure.

PlatformFormulaMonthly
Modal28,000 × $0.000694$19.43
RunPod Serverless7.78 hr × $2.72$21.16
Replicate, public model10,000 s × $0.0014 (setup and idle free)$14.00
Replicate, private model28,000 s × $0.0014$39.20
RunPod pod, always on730 hr × $1.59$1,160.70
Modal with 4 cores + 32 GiB28,000 s × ($0.000694 + $0.00012344)$22.89

Equal compute, execution only: if each platform billed only the 10,000 seconds of execution, Modal (GPU-only) would cost 10,000 × $0.000694 = $6.94, RunPod Serverless 2.78 hr × $2.72 = $7.56 and Replicate $14.00. Replicate's public-model row bills only that active time; Modal and RunPod bill container or worker time including start and idle, and Replicate private models bill setup, idle and active time (28,000 seconds, $39.20). The $14.00 public-model figure is therefore not directly comparable with the 28,000-second rows.

Workload B: steady, one A100 at 60% utilization

438 active hours in a 730-hour month.

PlatformZero overheadWith 10% overhead
Modal, GPU-only438 × $2.4984 = $1,094.30$1,203.73
Modal with 4 cores + 32 GiB438 × $2.9428 = $1,288.94$1,417.83
RunPod pod730 × $1.59 = $1,160.70$1,160.70
RunPod Serverless438 × $2.72 = $1,191.36$1,310.50
Replicate private, one instance online730 × $5.04 = $3,679.20$3,679.20

Break-even utilization = pod rate ÷ serverless rate: $1.59 ÷ $2.4984 = 63.6% (Modal, GPU-only) or 54.0% with the CPU and memory above; $1.59 ÷ $2.72 = 58.5% (RunPod Serverless). At 60% the pod is cheaper than RunPod Serverless but not than Modal until overhead is counted.

Workload C: two H100s at 80% utilization

1,168 active GPU-hours; 1,460 pod-hours.

PlatformFormulaMonthly
RunPod pods (PCIe)1,460 × $2.89$4,219.40
Modal, GPU-only1,168 × $3.9492$4,612.67 ($4,751.05 with 3% retry loss)
Modal with 8 cores + 64 GiB1,168 × $4.8380$5,650.75
RunPod pods (SXM)1,460 × $3.49$5,095.40
RunPod Serverless1,168 × $4.79$5,594.72
Replicate, active time1,168 × $5.49$6,412.32
Replicate, two deployments online1,460 × $5.49$8,015.40

H100 break-even utilization: Modal (GPU-only) versus a PCIe pod $2.89 ÷ $3.9492 = 73.2% (59.7% with CPU and memory); versus an SXM pod $3.49 ÷ $3.9492 = 88.4%; RunPod Serverless versus a PCIe pod $2.89 ÷ $4.79 = 60.3%. PCIe and SXM cards differ in interconnect, so the pod row is not a like-for-like comparison with every serverless H100.

What the advertised price misses

Idle hours are the main cost of a pod. Cold starts bill at the full per-second rate on RunPod Serverless, and model loading time depends on model size and platform configuration; no measured start time is assumed as a vendor fact. RunPod lists no data ingress or egress charge. Replicate's public-versus-private difference means a workload priced at $14 a month on a public model can cost $39 privately at low volume and $3,679 if kept warm.

Which workload fits

Spiky traffic and prototypes: Modal or RunPod Serverless, or Replicate public models where a suitable public model exists. Steady traffic near 60%: compare Modal and a pod using measured overhead. Sustained loads above roughly 60–75%: rented pods, sized to load.

Sources

Modal: modal.com/pricing, rates as read by Markaicode, 3 Sep 2026. RunPod: runpod.io/pricing (updated 13 Sep 2026), serverless billing docs; rates as verified by UsagePricing, 22 Sep 2026 and Omid Saffari, 25 Sep 2026. Replicate: replicate.com/pricing; hardware rates as reported by Markaicode, Aug 2026 and CheckThat.