NotCheapMAKE EVERY CREDIT COUNT

OpenAI Realtime API vs Gemini Live API vs Deepgram Voice Agent API: Real Voice AI Cost per 10,000 Minutes in 2026

The short answer: these three price live voice fundamentally differently, and the vendor with the "cheapest" headline rate depends entirely on which of several coexisting model tiers each one is quoting. OpenAI's Realtime API (gpt-realtime-2.1) bills in audio tokens — $32 per million input, $64 per million output — which at a typical 1,200 tokens per minute of audio works out to roughly $0.08–$0.12 per minute uncached, falling to $0.06–$0.08 with prompt caching; 10,000 minutes costs $1,152 uncached or about $773 well-cached, and an uncached, poorly-managed session can run far higher. Deepgram's Voice Agent API bundles speech-to-text, an LLM, and text-to-speech into one flat $4.50 an hour ($0.075/minute), or $750 for 10,000 minutes, telephony excluded. Google's Gemini Live pricing is the least settled of the three: its current native-audio-model rate ($3.00/M input, $12.00/M output for audio) works out to about $0.018/minute, or $180 for 10,000 minutes, but an older, still-referenced Gemini Live 2.5 Flash preview tier prices roughly 6.5 times cheaper still, and a separately listed "Gemini 3.8 Live" rate sits in between — three different numbers for what a buyer might reasonably call "the Gemini Live rate."

What each vendor bills

OpenAI Realtime API (OFFICIAL, OpenAI pricing documentation, July 2026). Billed in audio tokens, not minutes directly: gpt-realtime-2.1: $32 per million audio input tokens, $64 per million audio output tokens. Cached audio input drops to $0.40 per million — an 80-times reduction — making prompt caching the single largest lever on this API; getting it wrong means every turn re-bills the full system prompt at the uncached rate. A cheaper GPT-Realtime-Whisper tier is available at a flat $0.017 per minute for less demanding use cases, and a "mini" flagship variant runs roughly a third of the full model's per-minute cost. OpenAI has not published a single flat per-minute rate for the flagship model, consistent with its audio-token billing model; any per-minute figure is a derived approximation that depends on how much the assistant talks versus listens and how effectively the session caches its system prompt.

Gemini Live API (OFFICIAL for the rate itself, with a real multi-tier conflict across current documentation and third-party trackers). Google's current native-audio-model pricing (ai.google.dev): $0.50 per million text input tokens or $3.00 per million audio/video input tokens; $2.00 per million text output tokens or $12.00 per million audio output tokens (including thinking tokens). A separate, still-referenced Gemini Live 2.5 Flash preview tier is priced at $0.30 input / $2.00 output per million tokens — roughly a sixth of the native-audio rate above — and a third listing, for a model referred to as "Gemini 3.8 Live," shows $0.75 input / $4.50 output. These three rates cannot all be "the current Gemini Live price" simultaneously; they most likely reflect different model generations or preview-versus-stable status, but this article could not fully reconcile which one applies to a new integration built today, and flags this as a genuine open question to confirm directly with Google before budgeting. One additional documented quirk: silence during a continuously open audio stream bills at the same rate as active speech, confirmed directly by Google Cloud Billing support, because the model continuously processes the stream to detect speech regardless of whether anyone is talking.

Deepgram Voice Agent API (REPORTED, consistent across multiple 2026 pricing trackers). A single bundled rate covering speech-to-text, the LLM, and text-to-speech together: $4.50 per hour, or $0.075 per minute, reported as the most consistently cited figure for this product across independent sources. Telephony (a carrier like Twilio) is billed separately on top, typically reported around $0.014 per minute for US outbound calls, plus a small monthly phone-number fee.

Converting tokens to minutes

Formula: audio tokens per minute ≈ 1,200 (an ILLUSTRATIVE, commonly-cited approximation for continuous audio at typical bitrates, used consistently across the sources reviewed for OpenAI's Realtime API). This conversion is what turns OpenAI's and Gemini's per-token rates into a comparable per-minute figure; it is an approximation, not a vendor-published constant, and actual token consumption varies with audio quality settings and how much silence versus speech a session contains.

Cost per 10,000 minutes

Formula: cost = minutes × 1,200 tokens/minute × (input rate + output rate) ÷ 1,000,000, treating input and output as occurring symmetrically across a live session (both directions of audio flow continuously during a call).

Platform / tierPer-minute costCost per 10,000 minutes
Gemini Live 2.5 Flash preview (older tier)$0.00276$27.60
Gemini Live, current native-audio rate$0.018$180.00
Gemini 3.8 Live (separately listed rate)$0.0063$63.00
Deepgram Voice Agent API (bundled STT+LLM+TTS)$0.075$750.00
OpenAI Realtime, gpt-realtime-2.1, cached$0.077$772.80
OpenAI Realtime, gpt-realtime-2.1, uncached$0.115$1,152.00

The roughly 42-fold spread between Gemini's cheapest referenced tier ($27.60) and OpenAI's uncached flagship rate ($1,152) is a genuine, unresolved artifact of which specific model generation and caching state each figure represents — not a stable "Gemini is 42 times cheaper than OpenAI" conclusion a buyer should rely on without confirming which exact model and caching setup is actually quoted.

The caching trap on OpenAI

Formula: uncached penalty = (uncached rate − cached rate) ÷ cached rate. Moving from a well-cached to an uncached session raises OpenAI's per-minute cost from about $0.077 to $0.115 in this model — a 49% increase — but the real-world penalty can be far worse: one detailed cost breakdown reported that a poorly-cached, long-running session that re-bills its full system prompt on every turn can reach $0.46 per minute, roughly six times the well-cached rate. Because audio input caching depends on session design (keeping the same system prompt and conversation context structured so the API can reuse it), this is an engineering decision with direct, multiplying budget consequences, not a fixed vendor rate.

Silence billing on Gemini Live

Because Google confirmed that a continuously open audio stream bills for silence at the same rate as active speech, a voice agent that stays connected between turns (rather than closing and reopening the stream) accrues cost during dead air. Formula: wasted-silence cost = idle stream minutes × per-minute rate. For a support line with long hold times or slow customer responses, this can meaningfully inflate the effective per-conversation cost beyond what a simple "minutes actually spoken" estimate would suggest, and it applies regardless of which of Gemini Live's several rate tiers is in use.

Sensitivity

  1. Which Gemini Live rate applies. The single largest unresolved variable in this article; the cheapest and most expensive of the three current or recently current rates found in the sources reviewed differ by roughly 6.5x ($0.018 versus $0.00276 per minute).
  2. Caching effectiveness on OpenAI. A roughly 50% swing between well-cached and moderately-cached sessions, and up to 6x for poorly-cached long sessions.
  3. Silence and idle stream time on Gemini Live. Directly inflates cost on a per-open-connection basis, independent of actual speech content.
  4. Model tier choice on OpenAI. The "mini" variant and the flat-rate Whisper-based tier both cost meaningfully less than the flagship gpt-realtime-2.1 model, at some cost to conversational sophistication.

Budgeting traps

  • Assuming a demo session's cost predicts production cost. A short, well-behaved test session under-samples the caching and silence-billing effects that dominate a real, longer-running deployment.
  • Treating Gemini Live's cheapest referenced rate as the current, generally available one. Confirm directly which model generation your integration is actually billed against.
  • Ignoring telephony as a separate cost layer on Deepgram. The $0.075/minute figure covers STT, LLM and TTS only; a carrier like Twilio bills the phone connection itself separately.
  • Leaving an audio stream open during silence on Gemini Live. Closing and reopening the connection between turns, where the application architecture allows it, avoids paying for dead air at the same rate as speech.

What to ask before you buy

Ask Google directly which Gemini Live model generation and rate applies to a new integration started today, given the multiple current and recently current rates found in public documentation and trackers. Measure your own actual audio-token consumption on a realistic multi-turn session, including cache hit rate, before budgeting OpenAI Realtime API costs off the bare per-token rate card.