Stable Diffusion 3.5 vs FLUX vs Midjourney: Real Self-Hosted vs API Image Cost in 2026
The short answer: open weights are free, the GPU that runs them is not, and the crossover depends on volume more than on the model. At a 50% utilization of a rented RTX 4090, image generation costs about $4.60 per 1,000 images in GPU time alone and $6.30 once storage and maintenance are included. But a single always-on GPU is a fixed bill of about $694 a month, so at 1,000 images a month the effective cost is about $694 per 1,000 images, and at 10,000 images it is $69. Hosted FLUX APIs at $0.014 to $0.04 per image ($14–$40 per 1,000) beat a dedicated GPU until volume reaches roughly 17,000 to 50,000 images a month. Midjourney is a subscription benchmark, not a self-hosted model, and it has no official API.
What each route costs
Cloud GPU (OFFICIAL for the listed rows, REPORTED for the rest). RunPod's pricing page (July 2026) lists an RTX 4090 at $0.34 an hour, an A100 80GB at $1.19, an L40S at $0.79 and an H100 80GB from $1.99 an hour, on-demand with per-second billing. Other sources report higher secure-cloud rates: RTX 4090 $0.69, A100 SXM $1.49, H100 PCIe $2.89–$2.99 and H100 SXM $3.29–$3.49 an hour. Serverless workers bill only while active, at about $0.72 an hour for a 4090 worker, with cold starts billed. Network volume storage was reported at $0.07 per GB per month. The model below uses secure-cloud reference rates of $0.69, $1.49 and $2.99 an hour, which are ILLUSTRATIVE choices within the published range.
FLUX API (REPORTED). Black Forest Labs' API lists FLUX.2 [klein] from $0.014 per image, FLUX.2 [pro] from $0.03 (per one megapixel), FLUX.2 [flex] from $0.05 and FLUX.2 [max] from $0.07; FLUX1.1 [pro] at $0.04, [pro] Ultra at $0.06. Third-party hosts such as fal.ai bill FLUX.1 [schnell] at about $0.003 per megapixel and Stable Diffusion 3 at $0.035 per image. Prices rise with resolution. Stability's hosted Stable Image Core was listed at $0.03 per image. Stable Diffusion 3.5's own API credit rate was not captured (QUOTE-ONLY / UNKNOWN).
Licensing (REPORTED). FLUX.1 [schnell] is Apache 2.0 and allows commercial use; FLUX.1 [dev] is non-commercial; FLUX.2 [dev] is local-only, and commercial self-hosting of FLUX.2 models requires a license tier from Black Forest Labs (Builder for FLUX.2 [klein] at 10,000 images a month, Platform and Professional for 100,000, Enterprise custom), with tiers above Builder quoted by sales (QUOTE-ONLY / UNKNOWN). Stable Diffusion 3.5 weights are available for self-hosting under Stability AI's license, whose commercial terms were not reviewed here (QUOTE-ONLY / UNKNOWN).
Midjourney (REPORTED). Subscription plans of $10, $30, $60 and $120 a month buy fast GPU hours (3.3, 15, 30, 60), with unlimited Relax mode on Standard and above, extra fast time at $4 an hour and about 0.8 GPU minutes per standard image. Companies above $1 million in annual revenue must use Pro or Mega. No official API exists.
Model assumptions
Throughput per GPU is ILLUSTRATIVE for a 1-megapixel image with a full-size model: 300 images an hour on an RTX 4090, 700 on an A100 80GB and 1,200 on an H100. Fixed monthly costs are ILLUSTRATIVE: $7 storage for a 100 GB volume, about $83 of amortized setup (20 hours at $50 spread over 12 months) and $100 of maintenance (2 hours a month), a total of $190. A dedicated GPU runs 730 hours a month whether or not it is busy.
Formulas: GPU-only cost per 1,000 images = hourly rate ÷ (images per hour × utilization) × 1,000; dedicated monthly cost = rate × 730 + fixed; break-even monthly images = dedicated monthly cost ÷ API price per image.
Effective cost per 1,000 images by utilization
| GPU | Monthly cost incl. $190 fixed | GPU-only at 20% / 50% / 80% | With fixed costs at 20% / 50% / 80% |
|---|---|---|---|
| RTX 4090 ($0.69) | $694 | $11.50 / $4.60 / $2.88 | $15.84 / $6.34 / $3.96 |
| A100 80GB ($1.49) | $1,278 | $10.64 / $4.26 / $2.66 | $12.50 / $5.00 / $3.13 |
| H100 ($2.99) | $2,373 | $12.46 / $4.98 / $3.11 | $13.54 / $5.42 / $3.39 |
At high utilization the three GPUs converge near $3 to $4 per 1,000 images: the faster cards cost proportionally more per hour. Utilization is the variable that matters, not the GPU class.
Cost by monthly volume
| Images per month | Dedicated RTX 4090 (utilization) | Serverless 4090 worker | Dedicated H100 (utilization) | FLUX.2 [klein] API ($0.014) | FLUX.2 [pro] API ($0.03) | Midjourney |
|---|---|---|---|---|---|---|
| 1,000 | $694 per 1,000 (0.5%) | $142 per 1,000 | $2,373 per 1,000 (0.1%) | $14 | $30 | Standard $30 (13 fast hours), about $30 per 1,000 |
| 10,000 | $69 per 1,000 (4.6%) | $16.40 per 1,000 | $237 per 1,000 (1.1%) | $14 | $30 | Standard with Relax $30 (about $3 per 1,000), or 133 fast hours: Mega plus top-ups $413 (about $41 per 1,000) |
| 100,000 | $6.90 per 1,000 (45.7%) | $3.80 per 1,000 | $24 per 1,000 (11.4%) | $14 | $30 | Relax on Standard or higher; 1,333 fast hours would cost about $5,200 |
The serverless figure assumes $2.40 per 1,000 images of active time plus $140 a month of storage, setup and maintenance; cold-start time is not included. Midjourney's Relax mode trades queue time for a flat fee, and an "unlimited" plan is not an API for automated production.
Break-even against API
| Dedicated GPU | vs $0.014 per image | vs $0.03 | vs $0.04 |
|---|---|---|---|
| RTX 4090 | 49,550 images a month (22.6% utilization) | 23,123 (10.6%) | 17,342 (7.9%) |
| A100 80GB | 91,264 (17.9%) | 42,590 (8.3%) | 31,942 (6.3%) |
| H100 | 169,479 (19.3%) | 79,090 (9.0%) | 59,318 (6.8%) |
A serverless 4090 worker breaks even against a $0.014 API at about 12,000 images a month, against $0.03 at about 5,100, and against $0.04 at about 3,700. Against a $0.003 hosted FLUX.1 [schnell] endpoint, a dedicated 4090 would need more images than one card can produce in a month (105.6% utilization), so self-hosting does not beat that price on infrastructure cost alone.
Self-hosting starts to make sense at roughly 8% to 23% utilization against mid-priced APIs, and only above 100% of one card's capacity against the cheapest hosted schnell endpoints.
Sensitivity
- Throughput. Halving images per hour doubles the cost per 1,000 images at every utilization level; optimized inference can lift it several times.
- Idle time. A dedicated GPU that runs 24/7 at 5% utilization is 20 times more expensive per image than at 100%.
- Model license. Commercial self-hosting of FLUX.2 requires a paid tier; using schnell avoids that but at lower quality.
- Labor. Two hours of maintenance a month is $100; a team that needs monitoring, queueing and autoscaling spends far more.
- Resolution. API prices scale per megapixel, so 2K or 4K output can double or quadruple the API side of the comparison.
Budgeting traps
- Free weights, paid infrastructure. Storage, idle GPU time, engineering hours and licenses are the real cost.
- Cold-start billing. Serverless workers bill loading time.
- Comparing subscription "unlimited" with per-image APIs. Relax mode is queue time, not capacity.
- Ignoring rejects. Failed or discarded generations consume the same GPU time; divide by the usable rate for cost per usable image.
What to ask before you buy
Measure images per hour on your own model and resolution, forecast monthly volume, and compute break-even against the API price for that model. Ask any GPU host what happens to idle billing, storage and cold starts.