NotCheapMAKE EVERY CREDIT COUNT

GPT-Live-1 Pricing Explained: What a Real Voice AI Agent Costs in 2026

OpenAI shipped GPT-Live-1 in the API on September 10, 2026, at $0.05 per minute, billed per second. That number will show up in every headline about it, and it is real — but it is also the smallest and least interesting part of what a production voice agent actually costs. GPT-Live-1 is a full-duplex voice layer: it listens and speaks at the same time, handles interruptions, and holds the rhythm of a phone call. It does not, by design, do the thinking. Every real deployment pairs it with a separate backend model that reasons, calls tools, and decides what to say — and that backend is billed entirely separately, at that model's own token rate.

This article models the complete stack: voice minutes, backend tokens, tool calls, retries, and the gap between "cost per minute" and "cost per outcome that actually mattered." Where a number comes from OpenAI's own announcement or pricing page, we say so. Where we're building a scenario to make the math concrete, we label it as an assumption.

What's verified, directly from OpenAI

  • GPT-Live-1 became available in the API on September 10, 2026, after first shipping inside ChatGPT in July 2026.
  • Price: $0.05 per minute of conversation for the voice layer, billed per second (no rounding up to the next full minute).
  • It is full-duplex — it can listen while it is speaking, which is the architectural difference from a traditional speech-to-text → LLM → text-to-speech pipeline.
  • The voice layer natively supplies ASR transcripts and response text, supports keyword biasing (useful for order numbers, postcodes, product names), and can still expose explicit turn boundaries if your app wants them.
  • Telephony is supported, so the agent can answer or place phone calls, not just handle in-app voice.
  • Reasoning and tool use are delegated to a backend model of your choosing — "Responses delegation" uses an OpenAI Responses model; "client delegation" lets your application call any other model, agent, or service and decide what comes back.
  • On the consumer side (not the API), GPT-Live-1 is now the default ChatGPT Voice model for Go, Plus, and Pro; a lighter GPT-Live-1 mini is the default for Free. Voice there is bundled into the subscription with no separate meter — this article is about the API, where it's metered.
  • GPT-Live-1's predecessor architecture, GPT-Realtime-2.1 (a single model that handles speech, reasoning, and tool selection together), remains available; OpenAI has stated gpt-realtime is being retired January 20, 2027 in favor of it. GPT-Live-1 is a different design, not a replacement for Realtime in every case — it drops at least one input capability (image input) that Realtime-family models supported.

What's not fully settled: OpenAI's launch framing is architecture-first, and the promotional/standard-pricing distinction that has applied to some other 2026 model launches has not been publicly flagged for GPT-Live-1 as of this writing. Treat $0.05/minute as the current, not-yet-time-limited rate, but verify on OpenAI's live pricing page before locking a long-term budget — prices published at launch are prices at launch.

The three-layer cost stack

A GPT-Live-1 voice agent has at minimum three billable layers, and a production deployment usually has a fourth:

LayerWhat it coversWho bills itTypical rate
1. Voice (GPT-Live-1)Listening, speaking, turn-taking, interruption handlingOpenAI, per second$0.05/min ($3.00/hour)
2. Backend reasoningDeciding what to say, using tools, holding the conversation's logicOpenAI (or another provider), per tokenVaries by model — see below
3. Tool callsCalendar lookups, CRM writes, payment processing, order statusThe tool's own API/usage pricingVaries, often free at low volume
4. Infrastructure (if telephony)Inbound/outbound phone lines, SIP trunkingA telephony provider (e.g., Twilio-class services), not OpenAITypically $0.01–$0.02/min for calls, separate from OpenAI entirely

Layer 1 is fixed and precisely knowable. Layers 2–4 are not automatically the smaller cost — which layer dominates depends heavily on call duration, backend model choice, whether telephony and tools are involved, and how much human escalation the workload requires. In several of the scenarios below, the voice layer itself is the single largest direct cost; in others, backend tokens or telephony overtake it. The point isn't that one layer always wins — it's that the "$0.05/minute" headline describes only one of several genuinely significant cost components, and treating it as the whole picture undercounts the real bill in every scenario modeled here.

Backend model choices and what they cost

(Verified OpenAI API pricing as of September 2026; note that OpenAI has been cutting prices on this family through the year, so re-check before committing to a large deployment.)

ModelInput / output per million tokensBest fit for a voice agent
GPT-5.6 Luna$0.20 / $1.20Simple, high-volume, low-stakes turns — order status, FAQ-style answers
GPT-5.6 Terra$2.00 / $12.00General-purpose support and qualification conversations needing real reasoning
GPT-5.6 Sol~$4.00 / $20.00 (cut from $5/$30 in September 2026; OpenAI states this promotional rate holds at least into late November 2026)Complex objection-handling, sales conversations, anything where a wrong answer is costly
GPT-6 Astra$10.00 / $50.00Highest-stakes reasoning; the model OpenAI used in its own Agents API launch example

A voice agent doesn't need to pick one model for every call — routing simple turns to Luna and escalating only when the conversation gets complex is a legitimate cost-control strategy, and one we'll use in the scenarios below.

Estimating backend token cost per call

This is a NotCheapAI calculation built on labeled assumptions, not a vendor-published rate. Voice conversations don't map cleanly to token counts the way text chat does, because the backend model is typically re-fed a running transcript plus system instructions on every turn. A reasonable, conservative planning assumption for a moderately complex voice turn:

  • ~200–400 input tokens per turn (system prompt overhead is usually cached and cheap; the live transcript context is what varies)
  • ~60–150 output tokens per turn (a spoken response is short — nobody wants a 500-word answer read aloud)
  • A typical 3-minute call runs roughly 6–10 conversational turns

Using the midpoint of these ranges (300 input / 100 output tokens per turn, 8 turns) on GPT-5.6 Terra ($2/$12 per million tokens):

  • Input: 300 × 8 = 2,400 tokens → $0.0048
  • Output: 100 × 8 = 800 tokens → $0.0096
  • Backend cost for a 3-minute call: roughly $0.014

Compare that to the voice layer cost for the same call: 3 minutes × $0.05 = $0.15. At this assumption set, the voice layer is actually the larger cost — which flips the common intuition that "the LLM is the expensive part." That's specific to short, simple calls on a mid-tier backend model; it inverts quickly once you add tool calls, longer calls, or a pricier backend model, as the scenarios below show.

Worked scenarios

(All scenarios use labeled assumptions for call volume, duration, and backend model choice — verify against your own expected traffic before budgeting.)

Scenario 1: Small AI receptionist

Assumptions: 500 calls/month, average 2 minutes/call, simple routing and FAQ answers, GPT-5.6 Luna backend, no external tool calls, no telephony (in-app voice only).

Cost componentMonthly total
Voice (500 × 2 min × $0.05)$50.00
Backend tokens (500 calls × ~$0.006 avg, Luna rate)~$3.00
Total~$53/month

Cost per call: ~$0.11. At this scale, voice minutes dominate the bill by roughly 16 to 1 over backend reasoning.

Scenario 2: Customer-support agent

Assumptions: 2,000 calls/month, average 4 minutes/call, 1–2 tool calls per session (order lookup via an internal API, assumed free/internal), GPT-5.6 Terra backend for real reasoning about account issues.

Cost componentMonthly total
Voice (2,000 × 4 min × $0.05)$400.00
Backend tokens (2,000 calls × ~$0.019 avg, Terra rate, ~10 turns/call)~$38.00
Tool calls (internal API, assumed no marginal cost)$0
Total~$438/month

Cost per call: ~$0.22.

For other voice-agent billing models, compare Vapi, Retell, and Bland with ElevenLabs Voice Agent. For ticket-based support automation, see Gorgias AI Agent's per-resolution costs.

Scenario 3: Sales/lead-qualification agent

Assumptions: 800 calls/month, average 6 minutes/call (qualification conversations run longer), GPT-5.6 Sol backend for objection-handling quality, one external tool call per session to a CRM (assumed ~$0.01/call at typical CRM API pricing).

Cost componentMonthly total
Voice (800 × 6 min × $0.05)$240.00
Backend tokens (800 calls × ~$0.056 avg, Sol rate, ~12 turns/call)~$45.00
CRM tool calls (800 × $0.01)$8.00
Total~$293/month

Cost per call: ~$0.37.

Scenario 4: Higher-volume deployment with routing

Assumptions: 20,000 calls/month, average 3.5 minutes/call, routed backend (70% of calls stay on Luna for simple routing/FAQ, 30% escalate to Terra for anything requiring real reasoning), telephony required at an assumed $0.015/min for inbound lines (a third-party telephony cost, not OpenAI's).

Cost componentMonthly total
Voice (20,000 × 3.5 min × $0.05)$3,500.00
Backend tokens (14,000 Luna calls × ~$0.007 + 6,000 Terra calls × ~$0.017)~$200.00
Telephony (20,000 × 3.5 min × $0.015)$1,050.00
Total~$4,750/month

Cost per call: ~$0.24. Note that at this volume, telephony — a cost OpenAI has nothing to do with — is larger than the entire backend reasoning bill, and is frequently the line item founders forget to budget when they only look at OpenAI's pricing page.

Retries, failed calls, and human escalation

None of the scenarios above account for calls that don't go cleanly — and in voice specifically, failure modes are expensive in a way text chat isn't, because a failed voice turn often means re-explaining context, which burns both extra voice minutes and extra backend tokens simultaneously.

  • A misunderstood request that requires the caller to repeat themselves effectively adds 1–2 extra turns to a call — a small backend cost increase, but a proportionally larger voice-minute cost increase, since voice time is billed regardless of whether the extra turn was useful.
  • A call that needs human escalation doesn't stop your GPT-Live-1 bill from accruing during the handoff, and it adds a cost this article deliberately won't put a specific number on: human agent time. Outsourced and in-house support labor costs vary enormously by market, skill level, and channel, and inventing a precise blended rate here would be false precision dressed as data. What's defensible to say: if even 5–10% of calls in Scenario 2 require human escalation, and each escalation consumes several minutes of paid human time on top of the AI minutes already spent, the true fully-loaded cost of that segment of calls can materially or substantially exceed the AI inference cost, depending on the labor rate and how long the escalation runs. We're not assigning that a specific multiple, because the honest range is too wide to state responsibly. Model your own escalation rate and labor cost directly rather than borrowing a number from this article.
  • Failed or dropped calls (a caller hangs up, a tool call times out) still consume whatever voice minutes elapsed before the failure. At meaningful volume, a 5% "call didn't complete its purpose" rate effectively taxes your entire voice-minute line by the same 5%, since you paid for those minutes with nothing to show for them.

Why cost per successful outcome matters more than cost per minute

Every calculation above answers "what does this call cost," not "was this call worth it." Those are different questions, and conflating them is the most common mistake in voice-agent budgeting.

Take Scenario 3, the sales/qualification agent, at roughly $0.37/call. If that agent successfully qualifies and books a follow-up on 1 in 8 calls, the real unit economics aren't "$0.37 per call" — they're $0.37 × 8 = ~$2.96 per booked lead. Whether that's cheap or expensive depends entirely on what a qualified lead is worth to the business, a number this article can't supply because it's specific to your product and sales cycle, not to GPT-Live-1's pricing.

The same logic runs in the other direction for the receptionist scenario: at $0.11/call, an agent that successfully routes 95% of callers to the right department is extremely cheap per successful routing (~$0.116). An agent that only routes correctly 60% of the time, forcing repeat calls or human intervention on the rest, might be nominally "the same $0.11/call" while actually costing several times more per successfully resolved contact once you count the calls that had to happen twice.

The practical rule: track a completion or conversion metric specific to your use case — booked appointment, resolved ticket without escalation, qualified lead, correctly routed call — and divide your total monthly AI + telephony + escalation spend by that number, not by raw call count. Cost per minute is an input to that calculation, not a substitute for it.

Frequently asked questions

Is $0.05/minute the total cost of running a GPT-Live-1 voice agent? No. It's the cost of the voice layer only. Every real deployment adds a backend model's token cost on top, and most add tool-call costs, and many add telephony costs if the agent answers real phone lines.

Does GPT-Live-1 replace GPT-Realtime-2.1? Not universally. OpenAI has said gpt-realtime is being retired in January 2027 in favor of gpt-realtime-2.1, and separately launched GPT-Live-1 as an architecturally different product that splits conversation handling from backend reasoning. GPT-Live-1 also drops image input, a capability the Realtime family had.

Can I use a cheaper non-OpenAI model as the backend? Yes — "client delegation" lets you route to any model, agent, or service you choose, and decide what response comes back. This is exactly how a founder could pair GPT-Live-1's voice layer with a cheaper or more specialized backend model from another provider.

What's the single most commonly forgotten cost in a voice agent budget? Telephony, if the agent needs to answer or place real phone calls rather than handling in-app voice. It's billed by a separate provider, has nothing to do with OpenAI's pricing page, and at high call volumes (see Scenario 4) can exceed the entire AI reasoning cost.

How much does human escalation really add? This article deliberately doesn't assign it a number, because labor cost varies too widely by market and channel to state one responsibly. Model your own escalation rate and blended labor cost directly; treat any generic "$X per escalated call" figure you see elsewhere with real skepticism.

Hidden costs beyond the three-layer stack

Multi-language and accent handling. GPT-Live-1's launch materials cite a broader selection of voices across accents, dialects, and languages than its predecessor — a genuine capability improvement, but not a cost-free one if your backend model needs separate handling logic per language (different system prompts, different tool-call formatting for region-specific data like addresses or phone formats). This isn't a GPT-Live-1 line item; it's an engineering cost the voice layer doesn't eliminate.

Safety and content review overhead. A voice agent handling real customer conversations, especially in regulated contexts (healthcare scheduling, financial services, anything collecting payment information over the phone), typically needs additional guardrail logic layered on top of the base backend model — either through system-prompt constraints, a separate moderation pass, or human spot-checking of transcripts. None of this is priced by OpenAI; all of it is a real cost of running a compliant production agent, and it scales with call volume the same way support costs do.

Latency-driven abandonment. Independent testing cited around GPT-Live-1's launch put its turn latency at roughly 0.8 seconds in one comparison, against 1.4–1.6 seconds for older pipelined systems. That's a genuine improvement, but even a well-performing voice agent will lose some fraction of callers to hang-ups during any delay, tool-call wait, or misrouted turn — and every abandoned call still consumes the voice minutes that elapsed before the caller left. This is a variant of the "failed call" cost already discussed, worth calling out separately because latency (not just outright errors) is itself a cost driver most teams don't model until they see it in production.

Session and transcript storage. Full-duplex voice sessions generate transcripts and audio logs that many businesses are required to retain for quality assurance, dispute resolution, or compliance. Storage itself is usually cheap at reasonable volume, but the retrieval, redaction (removing payment details or other sensitive data before long-term storage), and review tooling around it is a real, if often small, addition to the total cost of running the agent — and one that scales with call volume and retention period, not with anything OpenAI bills directly.

Worked example: cost per resolved call

For Scenario 2 (the customer-support agent), assume a 65% successful-resolution rate: the share of calls that end without human escalation.

  • Base monthly cost (voice + backend + tools): ~$438, as calculated earlier
  • Assumed escalation rate: 35% of calls (700 of 2,000) require human handoff
  • If each escalation is separately tracked in the business's own support-cost system (deliberately not estimated here), the result is: cost per successfully AI-resolved call = $438 ÷ 1,300 ≈ $0.34, clearly distinct from the flat $0.22/call figure calculated earlier across all calls regardless of outcome
  • This is the number that should actually inform a "is this agent worth running" decision — not the blended per-call figure, and not the flat per-minute rate the industry tends to quote first