NotCheapMAKE EVERY CREDIT COUNT

Replicate vs fal.ai vs Modal: Real Serverless AI Inference Cost per 1 Million Runs in 2026

The short answer: all three bill by the second of GPU time, but at different rates on the same hardware, and Modal is consistently the cheapest of the three on identical GPUs — about 27–50% cheaper than Replicate (T4, A100, H100) and 12–37% cheaper than fal.ai (A100, H100) at the per-second rate level, with the exact gap varying by GPU class rather than holding at one flat percentage. On a medium workload (a 3-second image-generation task on an A100), 1 million runs costs about $4,200 on Replicate, $3,330 on fal.ai, and $2,082 on Modal. Replicate, now owned by Cloudflare following a November 2025 acquisition that closed in early 2026, has not changed its per-second rates since the deal, but it carries a real, separate cost risk the other two largely avoid: custom-model cold starts reported around 60 seconds, which are billed the same as active compute time.

What each vendor bills

Replicate (OFFICIAL, Replicate's pricing page). Per-second GPU billing, no published minimum increment: T4: $0.000225/sec ($0.81/hr). L40S: $0.000975/sec ($3.51/hr). A100 80GB: $0.001400/sec ($5.04/hr). H100: $0.001525/sec ($5.49/hr). Multi-GPU tiers scale roughly linearly (8x H100: $0.0122/sec, $43.92/hr). A separate "Official Models" catalog, launched January 2025, bills select models (including some third-party frontier LLMs resold through Replicate) on a per-token or per-output basis instead of raw GPU-seconds, which excludes cold-start overhead for those specific models. Cloudflare's acquisition of Replicate was announced in November 2025 and closed in early 2026; the per-second rate table has not moved since either event.

Modal (OFFICIAL, Modal's pricing documentation). Per-second GPU billing with no published minimum: T4: $0.000164/sec (~$0.59/hr). A100 80GB: $0.000694/sec ($2.50/hr). H100: $0.001097/sec ($3.95/hr). H200: $0.001261/sec ($4.54/hr). B200: $0.001736/sec ($6.25/hr). CPU is billed separately at $0.0000131 per physical core-second and memory at $0.00000222 per GiB-second for workloads that need dedicated CPU/RAM alongside the GPU. A $30/month free credit is available, and cold starts are reported under 2–4 seconds for typical containers — an order of magnitude faster than Replicate's custom-model cold-start figure.

fal.ai (REPORTED for per-second rates; OFFICIAL for the billing model). Primarily output-based pricing (per image, per second of generated video or audio) rather than raw GPU-seconds for its hosted model catalog, which fal.ai positions as more predictable for high-volume image generation specifically. For custom deployments billed by GPU-second, reported rates: A100 40GB: ~$0.00111/sec ($3.99/hr). H100 80GB: ~$0.00125/sec ($4.50/hr). A6000 48GB: ~$0.000575/sec ($2.07/hr). No permanent free tier; promotional credits are offered periodically rather than a standing allowance.

Three workload shapes and their per-run cost

All ILLUSTRATIVE. A: light task (a short classification or small-model inference call, 2 seconds on a T4-class GPU). B: medium task (a typical Stable-Diffusion-class image generation, 3 seconds on an A100-class GPU). C: heavy task (a short video-generation or large-model call, 30 seconds on an H100-class GPU).

Formula: cost per run = GPU per-second rate × seconds; cost per 1 million runs = cost per run × 1,000,000.

ShapeReplicateModalfal.ai
A: light (T4, 2s)$450$328N/A (no published T4 per-second rate)
B: medium (A100, 3s)$4,200$2,082$3,330
C: heavy (H100, 30s)$45,750$32,910$37,500

Modal is cheapest at every shape tested; on shape B, Modal is about 50% cheaper than Replicate and 37% cheaper than fal.ai on the same GPU class, purely from the underlying per-second rate difference — none of these figures yet account for cold starts, which change the real-world picture materially.

The cold-start trap

This is the single largest hidden cost difference between these three platforms, and it does not show up in any per-second rate table. Replicate's custom (non-Official-Models) deployments were reported with cold starts around 60 seconds, and that startup time is billed as active GPU-seconds on a per-second-metered platform — a bursty workload that triggers frequent cold starts on Replicate can pay for far more GPU-seconds than its actual compute work requires. Modal's cold starts were reported under 2–4 seconds, and fal.ai's hosted-model catalog is designed around pre-warmed inference engines specifically to avoid this problem for its output-priced models. Formula: cold-start tax = cold-start seconds × per-second rate × cold-start frequency. For a shape-B workload (A100, 3-second task) hit by a cold start on 10% of calls, Replicate's 60-second cold start adds roughly $0.084 per affected call ($0.0014/sec × 60s), which is 20 times the task's own $0.0042 compute cost — at 10% cold-start frequency across 1 million calls, that is an additional $8,400 on top of the $4,200 base compute figure, effectively tripling the real bill. Modal's 2–4 second cold start on the same math adds roughly 6.7% (2-second cold start) to 13.3% (4-second cold start) to average GPU time at the same 10% cold-start frequency on a 3-second task — a real addition, though still far smaller in absolute and proportional terms than Replicate's 60-second figure.

Replicate's Official Models: a different pricing track entirely

For models available through Replicate's Official Models catalog, pricing shifts to per-token (for LLMs) or per-output-unit (for image/video), excluding the cold-start-billing problem described above for those specific models. This is not comparable to the raw GPU-second figures in this article's tables — it is Replicate reselling a subset of models under a different metering scheme, closer to how a hosted API like OpenAI's or Together's prices its own catalog, and it should be evaluated separately from the per-second custom-model numbers.

Sensitivity

  1. Cold-start frequency. The dominant unmodeled cost driver on Replicate specifically; a workload with steady, high-frequency traffic that keeps containers warm avoids this almost entirely, while a bursty or low-traffic workload gets hit repeatedly.
  2. GPU class. Moving from A100 to H100 raises the per-second rate by roughly 9% (Replicate), 58% (Modal), or 13% (fal.ai) — Modal's H100 premium over A100 is proportionally the largest of the three, though its H100 rate is still the cheapest in absolute terms.
  3. Task duration. Linear in every platform's formula; a task that runs twice as long costs twice as much regardless of vendor.
  4. Output-based versus per-second pricing on fal.ai. For fal.ai's primary hosted-model catalog (not the custom-deployment rates used above), per-image or per-video-second pricing can be cheaper or more expensive than the equivalent GPU-second calculation depending on how efficiently fal.ai's inference engine packs each request.

Budgeting traps

  • Pricing a bursty workload off the per-second rate alone. Cold starts on Replicate specifically can multiply the effective cost several times over for low-frequency or spiky traffic.
  • Comparing Replicate's Official Models per-token pricing to another platform's per-second GPU pricing. These are two different metering schemes even within Replicate's own product line.
  • Assuming fal.ai's per-second custom-deployment rate applies to its hosted-model catalog. Most of fal.ai's traffic runs through output-based pricing, not the per-second figures in this article's tables.
  • Forgetting CPU and memory costs on Modal for workloads needing dedicated resources beyond the GPU. These are billed separately from the GPU-second rate.

What to ask before you buy

Ask Replicate directly for its current cold-start behavior on your specific model and expected traffic pattern, since a 60-second cold start at low request frequency can dominate total cost more than the GPU rate itself. Measure your own workload's typical task duration on each platform's actual hardware before comparing per-second rates, since GPU class and task length together determine the real per-run cost far more than which vendor's rate card looks lowest at a glance.