NotCheapMAKE EVERY CREDIT COUNT

Cartesia vs ElevenLabs vs OpenAI TTS: Real AI Voice Generation Cost per 10,000 Minutes in 2026

The short answer: these three price synthesis in three incompatible units — ElevenLabs and OpenAI's tts-1 family bill per character, Cartesia bills in credits per second of generated audio, and OpenAI's gpt-4o-mini-tts bills in tokens split between text input and audio output. Normalized to an ILLUSTRATIVE 750 characters per minute of narration, generating 10,000 minutes of audio costs roughly $112.50–$225 on OpenAI's tts-1 family, $150 on gpt-4o-mini-tts, and $375–$750 on ElevenLabs' API models, depending on quality tier. Cartesia cannot be given a clean 10,000-minute figure: its Scale plan includes 8,000,000 model credits a month, enough for about 8,889 minutes at 900 credits per minute, so 10,000 minutes needs roughly 1,000,000 credits (about 12.5%) more than the plan includes, and Cartesia does not publish a top-up or overage rate for credits beyond a plan's monthly allocation — the honest number is "$299 plus an unpriced overage," not a single dollar figure. ElevenLabs also sells a separate bundled Conversational AI product at $0.10 a minute, which is a different purchase from its raw TTS API and should not be confused with it. The characters-per-minute assumption drives every character-priced number in this article and should be replaced with your own measured speaking rate before budgeting.

What each vendor bills

ElevenLabs (REPORTED). Per-character pricing on its API: Multilingual v2/v3 at $0.10 per 1,000 characters ($100 per million); Flash at $0.05 per 1,000 ($50 per million). Separately, ElevenLabs cut its bundled Conversational AI product to $0.10 per minute in early 2026, a full voice-agent-minute rate distinct from the raw per-character TTS API price, and not directly comparable to it without knowing the underlying character count per minute. ElevenLabs' free tier is small (about 10,000 characters a month), and commercial use requires a paid plan.

Cartesia (OFFICIAL for the credit mechanism and plan allocations, REPORTED for the dollar-per-minute conversion; QUOTE-ONLY / UNKNOWN for any overage rate). Priced in credits, not characters: 15 credits per second of generated audio (900 credits per minute). Free: 20,000 credits/month (about 22 minutes), no commercial license. Pro: $5/month, 100,000 credits/month (about 111 minutes), adds commercial use and instant voice cloning. Startup: $49/month, 1,250,000 credits/month (about 1,389 minutes, roughly 23 hours), Pro voice cloning. Scale: $299/month, 8,000,000 credits/month (about 8,889 minutes, roughly 148 hours), higher concurrency. Enterprise: custom, contact sales. Each plan's credit allocation is a monthly cap, not a rate that can be extrapolated indefinitely: at the Startup allocation, the plan covers about 1,389 minutes a month for $49 flat; at Scale, about 8,889 minutes a month for $299 flat. A generation volume that exceeds a plan's included credits requires either upgrading to the next tier or paying for additional credits, and Cartesia does not publish a confirmed top-up or overage price per credit in the sources reviewed for this article; buyer guides note only that "the included model credits cover a demo or light integration test, not sustained production traffic" without stating what happens, in dollars, once that allocation is exhausted. This article therefore does not calculate a per-minute Cartesia rate by dividing a plan's monthly price by its included minutes and extrapolating that rate to higher volumes, because that rate is not what Cartesia would actually charge beyond the included allocation. Artificial Analysis separately normalized Cartesia's Sonic 3.5 model to roughly $49 per million characters as an independent cost estimate, a fifth-place price among tracked engines and about five times cheaper than ElevenLabs on a per-character basis, according to that source; this figure does not come from Cartesia's own plan pricing and is not used elsewhere in this article's tables.

OpenAI (OFFICIAL, OpenAI pricing documentation as read July 2026). tts-1: $15 per million characters. tts-1-hd: $30 per million. gpt-4o-mini-tts is token-based: $0.60 per million text input tokens plus $12 per million audio output tokens, which several independent sources converge on as roughly $0.015 per minute of generated audio at typical speech density. This is OpenAI's cheapest tier and the most straightforward to bundle with a model call already running on OpenAI's API.

Normalizing the unit: an explicit speech-speed assumption

This article uses 750 characters per minute of narration as its baseline conversion (ILLUSTRATIVE; actual density varies by language, punctuation, and speaking pace, commonly reported in the 600–900 character-per-minute range for English narration at a natural pace).

Formula: cost per minute = (characters per minute ÷ 1,000,000) × price per million characters, or the vendor's native credit/token formula where applicable.

Engine and tierCost per minute (at 750 chars/min)Cost per 10,000 minutes
OpenAI tts-1 ($15/1M chars)$0.01125$112.50
OpenAI gpt-4o-mini-tts (reported ~$0.015/min)$0.015$150.00
OpenAI tts-1-hd ($30/1M chars)$0.0225$225.00
ElevenLabs Flash ($50/1M chars)$0.0375$375.00
ElevenLabs Multilingual v2/v3 ($100/1M chars)$0.075$750.00
ElevenLabs Conversational AI (bundled, per-minute product)$0.10$1,000.00

Cartesia is not in this per-minute table because its plans are monthly credit caps, not a rate that holds at any volume. At 10,000 minutes a month (9,000,000 credits at 900 credits/minute), Cartesia's Scale plan ($299/month, 8,000,000 credits included) is about 1,000,000 credits (12.5%) short of covering the volume, and no confirmed top-up rate exists to price the shortfall. The honest figure for 10,000 minutes on Cartesia is $299 (the Scale plan floor) plus a QUOTE-ONLY / UNKNOWN overage, not a single extrapolated dollar amount.

Cost by monthly volume

Formula: cost = generated hours × 60 × cost per minute.

Engine100 hours (6,000 min)1,000 hours (60,000 min)10,000 hours (600,000 min)
OpenAI tts-1$67.50$675.00$6,750.00
OpenAI gpt-4o-mini-tts$90.00$900.00$9,000.00
OpenAI tts-1-hd$135.00$1,350.00$13,500.00
Cartesia Sonic, Scale plan$299.00 (flat; 5,400,000 of 8,000,000 included credits used)$299 + QUOTE-ONLY / UNKNOWN overage (54,000,000 credits needed, 46,000,000 over the 8,000,000 included)Enterprise tier likely required (540,000,000 credits needed); custom pricing, QUOTE-ONLY / UNKNOWN
ElevenLabs Flash$225.00$2,250.00$22,500.00
ElevenLabs Multilingual$450.00$4,500.00$45,000.00
ElevenLabs Conversational AI (bundled)$600.00$6,000.00$60,000.00

At 10,000 hours a month, OpenAI's cheapest tier costs about 15% of ElevenLabs Multilingual's price for raw synthesis, and ElevenLabs' bundled Conversational AI product (which includes more than raw TTS) costs nearly 9 times OpenAI's tts-1 rate for the equivalent minutes. Cartesia's Scale plan is the only volume in this table where its published price is a real, complete number: at 100 hours, the workload fits inside the plan's included credits, so $299 is the whole bill. At 1,000 and 10,000 hours, the workload requires several to dozens of times Scale's included allocation, and this article does not know, and will not guess, what that costs, because no public per-credit overage rate exists; a real deployment at this scale would need a quote from Cartesia's Enterprise tier.

Sensitivity to speech speed

Because every character-based price scales linearly with the characters-per-minute assumption, a faster or slower narration pace changes every figure in this article proportionally.

Characters per minuteElevenLabs Multilingual, cost per minuteCost per 10,000 minutes
600 (slower pace)$0.060$600.00
750 (baseline)$0.075$750.00
900 (faster pace)$0.090$900.00

A 20% faster speaking pace (750 to 900 characters per minute) raises every character-priced vendor's bill by exactly 20%. Cartesia's credit meter bills by generated audio duration (credits per second) and is genuinely unaffected by how dense the input text is. OpenAI's gpt-4o-mini-tts is less clean on this point: it bills a text-input-token charge ($0.60 per million) alongside the audio-output-token charge ($12 per million), so a denser or more verbose input script can raise its cost somewhat even at the same audio duration, though the audio-output charge (the larger of the two) still dominates and keeps the overall sensitivity to speech density far smaller than on a per-character API.

Commercial-use and cloning requirements

ElevenLabs' free tier does not permit commercial use. Cartesia's free tier similarly excludes commercial use and instant voice cloning, both unlocked starting at its $5-a-month Pro tier. OpenAI's pay-as-you-go pricing carries no separate commercial-use gate reported in the sources reviewed. Professional or instant voice cloning is a paid-tier feature on both ElevenLabs and Cartesia; OpenAI's TTS models do not offer voice cloning from a reference sample in the same way.

Sensitivity

  1. Which conversion rate is real. ElevenLabs' own credit-per-character schedule versus published per-character rates can differ by plan tier; the API rate cited here is the closest to a stable per-character reference.
  2. Model quality tier. Moving from OpenAI's tts-1 to tts-1-hd exactly doubles the per-minute cost for a quality upgrade; ElevenLabs' Flash-to-Multilingual jump is exactly 2x for the same reason.
  3. Bundled versus raw API pricing. ElevenLabs' $0.10-per-minute Conversational AI figure includes more infrastructure than a raw synthesis call and should never be compared directly to a per-character API rate without adjustment.
  4. Which side of Cartesia's included-credit line you land on. Inside the allocation, Cartesia's plan fee is a flat, fully known number ($49 for about 1,389 minutes on Startup, $299 for about 8,889 minutes on Scale); the moment a workload crosses that line, the real cost becomes unknown, since no public overage rate exists. This is a much larger swing than the difference between vendors at volumes that stay within a plan's included credits.

Budgeting traps

  • Comparing a bundled voice-agent rate to a raw synthesis rate. ElevenLabs' Conversational AI product and its Multilingual API model are not the same purchase.
  • Ignoring the characters-per-minute assumption. Every character-priced vendor's total scales directly with this number, which the article's own baseline explicitly flags as ILLUSTRATIVE.
  • Treating a plan's included-credit rate as a price that scales. Cartesia's $49 and $299 plans are flat fees for a fixed monthly credit allocation, not a per-minute rate you can multiply out to any volume; once usage exceeds the included credits, the real additional cost is not publicly known.
  • Assuming the free tier supports production use. Commercial use is gated behind paid tiers on both ElevenLabs and Cartesia.

What to ask before you buy

Measure your own actual characters-per-minute or seconds-of-audio-per-turn on a sample of real scripts, then convert every vendor's rate card into the same per-minute figure using that measured rate rather than an assumed round number. Ask specifically whether a quoted rate is for the raw synthesis API or a bundled voice-agent product.


ARTICLE 14