NotCheapMAKE EVERY CREDIT COUNT

ElevenLabs Voice Agent: Real Production Cost and Overage Economics (2026)

Short answer: ElevenLabs Agents are billed per conversation minute against a plan-specific included-minutes bucket, not against the same shared credit pool used for text-to-speech and other ElevenLabs products. Every plan from Free through Business includes a minutes allowance and a concurrency cap; going over either costs extra — $0.08/min for minutes beyond your bucket, and double that ($0.16/min) if you exceed your plan's concurrent-call limit during a burst. None of this includes the LLM powering the conversation, which is billed separately based on the model you choose. For steady, moderate production volume within your plan's included minutes, ElevenLabs' effective per-minute cost is close to its advertised $0.08–$0.10/min. For bursty or under-provisioned production traffic, the concurrency penalty and LLM pass-through can push the real cost noticeably higher than the plan sticker price suggests.

Pricing at a glance

PlanMonthly priceIncluded call minutesConcurrent calls included
Free$0154
Starter$6756
Creator$22 ($11 first month)27510
Pro$991,23820
Scale$2993,73830
Business$99012,37540
EnterpriseCustomCustomCustom
Usage beyond plan limitsRate
Additional minutes (within concurrency limit)$0.08/min
Burst minutes (over the concurrency limit)$0.16/min (double the standard overage rate)
Text messages sent by the agent$0.003 each
LLM (the model powering the conversation)Billed separately by ElevenLabs, at the connected model's own per-token rate — not included in the per-minute Agents price

Source: elevenlabs.io/pricing/agents and elevenlabs.io/pricing.

What the per-minute rate does — and doesn't — cover

The $0.08–$0.10/min Agents rate covers ElevenLabs' own voice pipeline: real-time speech-to-text, text-to-speech generation, and the call orchestration layer for a conversational agent. It does not cover:

  • The LLM. Every Agents call passes through to whichever language model you've configured (GPT-4o, Claude, or others), and that model's token cost is invoiced separately on top of the per-minute Agents charge. Depending on the model and system-prompt length, this has been reported at roughly $0.02–$0.05 per minute of conversation for a mid-sized system prompt on a model like GPT-4o — a cost that rises further with larger prompts or a more expensive frontier model.
  • Telephony. Connecting the agent to an actual phone number requires a separate carrier (e.g., Twilio), billed outside ElevenLabs entirely.
  • Anything drawing from the separate shared credit pool — standard text-to-speech generation, transcription, dubbing, or voice design — which is billed under ElevenLabs' character/credit system, not the Agents per-minute system, even if used to prepare assets for the same voice agent.

Hidden costs to model before you buy

  • Burst pricing doubles your rate, not just adds a small premium. Exceeding your plan's concurrency cap — even briefly, during a traffic spike — bills every minute over that threshold at $0.16/min instead of $0.08/min. A plan sized for average load rather than peak load will systematically understate your real bill.
  • The LLM is the most commonly underestimated line item. Because the Agents price is quoted per minute and looks complete, it's easy to forget the LLM invoice arrives separately. At an estimated $0.02–$0.05/min on a capable model, LLM cost alone can add 25–60% on top of the advertised Agents rate before any overage is even considered.
  • Moving up a plan tier to avoid overage doesn't always pay off immediately. Pro's 1,238 included minutes cost $99/month (about $0.08/min if fully used); Scale's 3,738 minutes cost $299/month (also about $0.08/min). If your actual usage sits well below a tier's included bucket, you're effectively paying the same or a higher blended rate as staying on the lower tier and paying overage for the difference — model your actual expected minutes against each tier's bucket before upgrading preemptively.
  • The Free and Starter plans' small concurrency caps (4 and 6 calls) make them unsuitable for any real production traffic — even a modest simultaneous-call spike will trigger burst pricing immediately.
  • Text messages are billed per message, separately from voice minutes — relevant for agents that mix voice and SMS/chat channels.

Worked scenarios

We model a support/scheduling voice agent using GPT-4o as the LLM (estimated $0.035/min mid-range) at three monthly minute volumes, choosing the plan tier that best fits each volume.

Scenario A — 1,000 minutes/month, steady load (no bursts)

  • Best-fit plan: Pro ($99/month, 1,238 minutes included — covers this volume with no overage)
  • LLM cost: 1,000 × $0.035 = $35
  • Total: $99 + $35 = $134/month → $0.134 per minute (all-in)

Scenario B — 5,000 minutes/month, steady load (no bursts)

  • Best-fit plan: Scale ($299/month, 3,738 minutes included; 1,262 minutes overage)
  • Overage: 1,262 × $0.08 = $100.96
  • LLM cost: 5,000 × $0.035 = $175
  • Total: $299 + $100.96 + $175 = $574.96/month → $0.115 per minute (all-in)

Scenario C — 15,000 minutes/month, with a traffic pattern that regularly exceeds Business's 40-concurrency cap during peak hours (assume 10% of minutes are burst minutes)

  • Best-fit plan: Business ($990/month, 12,375 minutes included; 2,625 minutes overage)
  • Of the 2,625 overage minutes, assume 10% of total volume (1,500 minutes) qualifies as burst-rate rather than standard overage, and the remaining 1,125 overage minutes bill at the standard rate
  • Burst-rate minutes: 1,500 × $0.16 = $240
  • Standard overage minutes: 1,125 × $0.08 = $90
  • LLM cost: 15,000 × $0.035 = $525
  • Total: $990 + $240 + $90 + $525 = $1,845/month → $0.123 per minute (all-in)

Break-even: ElevenLabs vs. a modular stack (Vapi/Retell-style)

Comparing Scenario B ($574.96/month all-in on ElevenLabs, including LLM) against a modular Retell stack using the same GPT-4o LLM assumption used throughout this article ($0.035/min): Retell's own published component rates are $0.055/min (voice infrastructure) + $0.015/min (standard TTS) + $0.015/min (telephony), and swapping in GPT-4o at $0.035/min in place of a lightweight LLM gives $0.055 + $0.015 + $0.015 + $0.035 = $0.120/min. At 5,000 minutes, that's 5,000 × $0.120 = $600/month. Against ElevenLabs' $574.96/month, that's a $25.04/month difference. At this volume, ElevenLabs' bundled per-minute agent pricing is modestly cheaper than a comparable modular Retell stack running the same LLM — the gap is narrow enough that it can flip with small changes to either platform's rates, but it's consistent with ElevenLabs' own TTS being bundled into a lower blended rate at Scale-tier volume than buying comparable TTS quality separately would cost on a modular platform.

The comparison shifts toward modular platforms (Vapi, Retell) at very high volume with heavy customization needs, where their at-cost, no-markup component pricing and ability to swap in cheaper LLMs or open-source models on a per-call basis gives more room to optimize than ElevenLabs' fixed per-minute Agents rate.

Who pays more — and when the answer flips

For steady, plan-appropriate volume, ElevenLabs' effective per-minute cost (including a typical LLM) lands in the $0.11–$0.13/min range across the scenarios modeled — competitive with, and in the mid-volume case cheaper than, a comparably-configured modular voice stack.

The economics turn unfavorable specifically when traffic is bursty relative to your plan's concurrency cap. A team on Pro (20 concurrent calls) running a marketing campaign that spikes to 35 simultaneous calls will see a meaningful share of that spike billed at the $0.16/min burst rate - in that situation, right-sizing concurrency (moving up a tier before a known traffic spike, not after) is the single highest-leverage cost lever available on this platform.

If you are comparing managed minutes with an assembled stack, see Vapi vs Retell vs Bland; for non-agent speech generation, compare ElevenLabs API cost per minute.

Limitations and uncertainty

  • The $0.02–$0.05/min LLM cost range used in the worked scenarios is an estimate based on typical system-prompt sizes on GPT-4o-class models; your actual LLM cost depends entirely on prompt length, model choice, and conversation turn count, and should be modeled against your specific configuration and the LLM provider's own current token pricing.
  • Scenario C's assumption that 10% of total minutes qualify for burst-rate billing is illustrative, meant to show the mechanism — your actual burst exposure depends on your specific traffic pattern and requires monitoring actual concurrent-call peaks against your plan's cap.
  • The Retell comparison in the break-even section uses a $0.120/min figure derived from Retell's published component rates ($0.055/min voice infrastructure, $0.015/min standard TTS, $0.015/min telephony) using the same GPT-4o LLM assumption ($0.035/min) applied elsewhere in this article, in place of Retell's own lightweight-LLM default — not an official Retell rate for this exact bundle.
  • ElevenLabs has changed Agents pricing during 2026, including a reported reduction in the per-minute Agents rate; reconfirm current rates directly against elevenlabs.io/pricing/agents before finalizing a budget, since usage-based AI pricing in this category has moved more than once within the year across every vendor we've covered.

Official sources

  • elevenlabs.io/pricing
  • elevenlabs.io/pricing/agents
  • elevenlabs.io/pricing/api (for LLM and adjacent product rates referenced for context)