NotCheapMAKE EVERY CREDIT COUNT

ElevenLabs vs Cartesia vs OpenAI Text-to-Speech: Real TTS API Cost at 1 Billion Characters in 2026

The short answer: OpenAI deprecated its entire character-priced TTS line — tts-1, tts-1-hd and the gpt-4o-mini-tts snapshots — on October 1, 2026, with removal scheduled for January 6, 2027, replacing them with gpt-realtime-2.1-mini, a token/audio-priced model with no published character-rate equivalent. The deprecated models remain callable during the transition window, and at 1 billion input characters their reported rates were $15,000 (tts-1), $30,000 (tts-1-hd), and an ILLUSTRATIVE $16,000–$17,000 character-equivalent for gpt-4o-mini-tts — all now transitional figures, not OpenAI's current recommended pricing. Cartesia's Scale-tier rate works out to roughly $37,000, ElevenLabs' fast models (Flash and Turbo v2.5) cost about $50,000, and ElevenLabs' highest-quality model (Multilingual v2 or Eleven v3) is the most expensive at about $100,000. OpenAI's current recommended model, gpt-realtime-2.1-mini, cannot be fairly placed in this 1-billion-character comparison at all without inventing an unsupported character-to-token conversion, since OpenAI prices it by audio tokens, not characters.

Billing units differ heavily here, exactly as this category warns

ElevenLabs bills in credits that convert to characters by model: roughly 1 credit per character on its highest-quality Multilingual v2/Eleven v3 models, and about 0.5 credits per character on the faster Flash and Turbo v2.5 models — so the same subscription credit balance produces roughly twice as many characters on Flash as on Multilingual. Cartesia also bills through credit-based plans (Starter, Scale, Enterprise) rather than one flat per-character number; a Scale-tier blended rate is reported around $37 per million characters, but plan allowances should be confirmed directly. OpenAI bills two different ways across its own TTS lineup: tts-1 and tts-1-hd are priced per character (as a per-1M-character rate), while gpt-4o-mini-tts is priced per token, separately for text input and audio output. None of these is "per character of finished speech" in a single, comparable sense, and converting tokens to characters requires an explicit, stated assumption.

What each vendor actually bills

ElevenLabs (OFFICIAL, elevenlabs.io pricing, as of July 2026). Flash v2.5 and Turbo v2.5: $50 per 1 million characters (about 0.5 credits per character on subscription plans). Multilingual v2 and Eleven v3: $100 per 1 million characters (about 1 credit per character). Professional voice cloning: roughly 1.5x the base model's credit cost. Subscription tiers bundle credits at a discount to pay-as-you-go: Creator $22/month (121,000 credits), Pro $99/month (600,000 credits), Scale $299/month (1.8 million credits), Business $990/month, with API overage beyond a plan's credits commonly reported around $0.10 per 1,000 characters (equivalent to the standard $100/1M rate). Flash v2.5 is the real-time path, reported at roughly 75ms time-to-first-audio over streaming; Multilingual v2 and Eleven v3 trade that latency for broader language coverage (29 languages on Multilingual v2, 70-plus on Eleven v3) and higher-fidelity prosody.

Cartesia (REPORTED; Cartesia publishes credit-based plans rather than one published flat per-character rate, so this figure should be confirmed directly). Sonic (and the newer Sonic 3.5) is built for low-latency conversational use, reported sub-90ms streaming time-to-first-audio and 42-language coverage on Sonic 3.5. At a reported 1 credit per character on the Scale plan, the effective rate works out to about $37 per million characters; professional voice cloning is reported at about 1.5 credits per character, consistent with ElevenLabs' cloning premium structure. Cartesia's plan allowances (Free, Starter, Scale, Enterprise) and their exact character ceilings are not fully itemized on its public pricing page.

OpenAI (OFFICIAL model names and dates, current as of October 2, 2026). Deprecated October 1, 2026, with removal scheduled for January 6, 2027: tts-1, tts-1-hd, and the gpt-4o-mini-tts snapshots. These remain callable through the transition window at their prior rates — tts-1: $15 per 1 million characters. tts-1-hd: $30 per 1 million characters. gpt-4o-mini-tts: token-based, $0.60 per 1 million text-input tokens plus $12 per 1 million audio-output tokens — but should be treated as a transitional comparison point for existing integrations, not a clean long-term choice for new ones. OpenAI's recommended replacement is gpt-realtime-2.1-mini, which bills by audio/token consumption rather than input characters; OpenAI does not publish a character-based rate for it, and this article does not invent one. For live, low-latency conversation, GPT-Realtime models price audio at $32/$64 per 1 million input/output tokens, and a dedicated GPT-Realtime-Translate line is reported at $0.034 per minute — both separate from the deprecated batch TTS models above.

Normalizing the workload

All ILLUSTRATIVE. 1 billion input characters, standard commercial (non-cloned) voice, English baseline, low-latency API use where each vendor's real-time-capable model is used. To express this as spoken audio, this article uses a stated conversion of 900 characters per minute of speech (roughly 150 words per minute), giving about 18,519 hours of output audio; a 750-characters-per-minute assumption (used in one sourced comparison) would instead give about 22,222 hours. Neither figure is a vendor-published constant — actual characters-per-minute varies with language, punctuation density and speaking rate — so any hour-based comparison carries this conversion's uncertainty on top of the dollar comparison.

Cost at the normalized workload

Formula: cost = characters ÷ 1,000,000 × rate per million, with OpenAI's gpt-4o-mini-tts converted from its token rate using the 900-characters-per-minute assumption and a further ILLUSTRATIVE estimate of tokens per character (this conversion is explicitly a sensitivity range, not a precise figure, since OpenAI's tokenizer does not map one-to-one with characters).

Provider / modelRate per 1M charactersCost at 1 billion characters
OpenAI tts-1 (DEPRECATED Oct 1, 2026; removal Jan 6, 2027)$15$15,000
OpenAI gpt-4o-mini-tts (DEPRECATED Oct 1, 2026; ILLUSTRATIVE character-equivalent, two independent estimates)~$16–$17~$16,000–$16,667
OpenAI tts-1-hd (DEPRECATED Oct 1, 2026; removal Jan 6, 2027)$30$30,000
Cartesia Sonic, Scale tier (REPORTED; confirm directly)~$37~$37,000
ElevenLabs Flash v2.5 / Turbo v2.5$50$50,000
ElevenLabs Multilingual v2 / Eleven v3$100$100,000
OpenAI gpt-realtime-2.1-mini (current recommended replacement)Not publishable — priced by audio/token consumption, not charactersCannot be normalized to 1 billion characters without an unsupported conversion; see below

Professional voice cloning raises the two credit-based vendors' totals by about 50%: ElevenLabs Multilingual with cloning reaches roughly $150,000, and Cartesia's Scale-tier rate with cloning reaches roughly $55,500. Standard (non-cloned, library) voices are assumed throughout the main table.

Why OpenAI's current recommended model cannot be placed in this table

gpt-realtime-2.1-mini is priced by audio and token consumption, the same billing shape as OpenAI's GPT-Realtime line, not by input characters the way tts-1, tts-1-hd and gpt-4o-mini-tts were. Converting that into a "$ per 1 million characters" figure would require assuming both a characters-per-token ratio and an audio-tokens-per-second-of-speech ratio — two separate, non-trivial conversions stacked on top of each other, each of which OpenAI does not publish as a fixed constant. Earlier in this article, a single characters-per-minute assumption was already flagged as a source of roughly 20% uncertainty on its own; stacking a second unsupported conversion on top of that to price a token-based model would produce a number precise-looking enough to be mistaken for a quote, while actually resting on two stacked guesses. The honest comparison point for gpt-realtime-2.1-mini is a workload defined in audio minutes or tokens directly, not 1 billion input characters; a reader needing that comparison should convert their own actual script length to OpenAI's published token/audio rate using their own content's measured tokenization, not a borrowed industry-wide ratio.

Why the spread is this wide

The gap is not primarily about raw audio quality; it tracks what each price point is built for. OpenAI's tts-1 and gpt-4o-mini-tts are bundled, general-purpose options living inside a platform most AI teams already pay for, with no dedicated low-latency conversational product beyond the separately priced Realtime line. Cartesia and ElevenLabs' Flash and Turbo tiers are purpose-built for real-time conversational agents, where sub-100ms time-to-first-audio has engineering value beyond the per-character rate. ElevenLabs' Multilingual v2 and Eleven v3 sit at the top of the range because they are priced for long-form, high-fidelity, multilingual narration rather than utility voice — a 2x premium over its own faster models for the same vendor, which is a larger gap than the difference between vendors at the fast end.

Sensitivity

  1. Which ElevenLabs or Cartesia model tier is used. A 2x swing within ElevenLabs alone (Flash/Turbo versus Multilingual/Eleven v3) is larger than the gap between several different vendors at the fast end.
  2. Voice cloning. Adds roughly 50% to the credit cost on both ElevenLabs and Cartesia; standard library voices avoid this entirely.
  3. The characters-per-minute conversion, when expressing cost per hour of audio. A 750 versus 900 characters-per-minute assumption changes the implied hours by about 20% for the same dollar total.
  4. Subscription credits versus pay-as-you-go overage. ElevenLabs' bundled subscription credits are cheaper per character than its reported ~$0.10/1,000-character overage rate; budgeting off the wrong one misstates cost at volume.
  5. Whether a new integration lands on OpenAI's deprecated character-priced models or its current gpt-realtime-2.1-mini. The former have a firm removal date (January 6, 2027) and a different billing basis entirely from the latter, so this choice affects both cost and long-term integration stability.

Budgeting traps

  • Comparing ElevenLabs' or Cartesia's credit price to OpenAI's per-character price without converting credits to characters first. The credit-to-character ratio differs by model within each vendor.
  • Using OpenAI's token-based gpt-4o-mini-tts or gpt-realtime-2.1-mini rate as if it were directly comparable to a per-character rate. Any such comparison requires an explicit, stated characters-per-token assumption; this article treats the gpt-4o-mini-tts conversion as an illustrative range and declines to invent one at all for gpt-realtime-2.1-mini.
  • Budgeting a new OpenAI TTS integration against tts-1, tts-1-hd or gpt-4o-mini-tts. All three were deprecated October 1, 2026, with removal scheduled for January 6, 2027; the recommended current model is gpt-realtime-2.1-mini, priced on an entirely different basis.
  • Pricing a production voice agent off ElevenLabs' or Cartesia's cheapest published number. The low-latency, agent-ready tiers (Flash v2.5, Sonic) are not the same models as the budget comparison point, and plan allowances should be confirmed directly.
  • Forgetting the cloning premium. A roughly 50% uplift applies on both ElevenLabs and Cartesia for professional voice cloning, not included in standard per-character rates.
  • Assuming Cartesia's $37/1M figure is a published flat rate. It is a reported effective rate from one plan tier, not a rate Cartesia states as a single number on its pricing page.

What to ask before you buy

Confirm Cartesia's current plan allowances and effective per-character rate directly, since it does not publish one flat number the way OpenAI and ElevenLabs do. Ask OpenAI for the actual token count your typical text produces before relying on any characters-per-token conversion for gpt-4o-mini-tts or gpt-realtime-2.1-mini, and measure your own characters-per-minute ratio on representative scripts before converting any of these dollar figures into a cost-per-hour-of-audio number. Before committing to a new integration, confirm gpt-realtime-2.1-mini's current rate and billing shape directly against OpenAI's pricing page, since the deprecated character-priced models in this table will not be available past January 6, 2027.