NotCheapMAKE EVERY CREDIT COUNT

Deepgram vs AssemblyAI vs OpenAI Speech-to-Text: Real Transcription API Cost per 10,000 Hours in 2026

The short answer: for straightforward batch English transcription, AssemblyAI's Universal-2 model is the cheapest of the three at about $1,500 for 10,000 hours ($0.15/hour), OpenAI's current recommended model (gpt-transcribe) costs about $2,700, and Deepgram's Nova-3 batch rate is about $2,580. Once speaker diarization is added, the order changes less than the dollar amounts: Deepgram's diarization add-on is included in its committed Growth tier but costs an extra $0.0044 a minute (about $2,640 for 10,000 hours) on pay-as-you-go, pushing a feature-complete Deepgram job to about $5,220. For real-time streaming, which all three price well above batch, Deepgram Nova-3 streaming runs about $4,620, AssemblyAI's Universal-3.5 Pro Realtime about $4,500, and OpenAI's current realtime model (gpt-live-transcribe) costs about $10,200. As of October 2026, OpenAI's recommended default transcription path is gpt-transcribe for batch and gpt-live-transcribe for realtime; the older whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize models were deprecated on August 26, 2026 and are scheduled for shutdown on February 26, 2027 — they may still be callable during the transition window but are not OpenAI's current recommended choice for new integrations.

The billing-unit mismatch, up front

All three vendors price per minute of audio, which looks like a clean, comparable unit, but what that minute buys differs. Deepgram bills per second of audio processed (not rounded to the minute), with diarization, redaction, keyterm prompting and entity detection billed as separate per-minute add-ons, and its Audio Intelligence features (sentiment, topic detection, summarization) billed per token, not per minute. AssemblyAI publishes an hourly reference rate but bills per minute, with speaker diarization reported bundled at no extra per-minute charge on its core model and separate, usage-based pricing for LLM-based features like summarization. OpenAI bills transcription as a token-metered API call in practice, but publishes an effective per-minute rate for its audio models; it offers no native diarization on its flagship transcription model, and a "diarize" variant is a separate, differently priced path. None of the three rounds audio the same way, so very short clips are billed somewhat differently across vendors, with Deepgram's per-second billing avoiding the rounding penalty the others can carry on short clips.

What each vendor actually bills

Deepgram (OFFICIAL, deepgram.com/pricing, pay-as-you-go). Nova-3 pre-recorded (batch), monolingual: about $0.0043 per minute (~$0.26/hour). Nova-3 streaming, monolingual: about $0.0077 per minute (~$0.46/hour), though other sourced figures for streaming range from $0.0059 to $0.0112 depending on language mode and whether diarization is bundled; this article uses $0.0077 as the base streaming rate and prices diarization separately. Multilingual streaming: about $0.0058–$0.0092 per minute depending on source. Speaker diarization: an additional reported $0.0044 per minute on pay-as-you-go (lower, and sometimes bundled, on committed Growth pricing). Growth tier: $4,000-plus a year prepaid, with discounts up to 20% off pay-as-you-go rates. A $200 free credit does not expire. Audio Intelligence features price per 1,000 tokens ($0.0003 input, $0.0006 output on pay-as-you-go) rather than per minute.

AssemblyAI (OFFICIAL structure; per-minute rates REPORTED from AssemblyAI's published hourly figures). Universal-2 (core batch model): about $0.15 an hour ($0.0025 a minute) — the cheapest batch rate found among the three. Universal-3.5 Pro, async: about $0.21 an hour ($0.0035 a minute), reported as AssemblyAI's highest-accuracy batch model (4.5% WER on AssemblyAI's own June 2026 benchmark across 80,000-plus files). Universal-3.5 Pro Realtime (streaming): about $0.45 an hour ($0.0075 a minute), separately, Universal-Streaming reported at about $0.15 an hour for a lower-accuracy real-time tier. A $50 free credit is offered. Summarization and other LLM-based add-ons bill separately and are not included in the per-minute figures above.

OpenAI (OFFICIAL model names and rates, current as of October 2, 2026). gpt-transcribe, OpenAI's current recommended/default path for file (batch) transcription: $0.0045 per minute (~$0.27/hour). gpt-live-transcribe, OpenAI's current realtime transcription model: $0.017 per minute (~$1.02/hour) — a dedicated, separately priced realtime path, not a batch model adapted for streaming. Deprecated as of August 26, 2026, with shutdown scheduled for February 26, 2027: whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize. These deprecated models remain callable during the transition window — gpt-4o-transcribe at a reported $0.006 per minute and gpt-4o-mini-transcribe at a reported $0.003 per minute — but should be treated as a transitional option for existing integrations, not OpenAI's recommended starting point for new ones. Diarization support on the current gpt-transcribe/gpt-live-transcribe line should be confirmed directly against OpenAI's current documentation.

Normalizing the workload

All ILLUSTRATIVE. 10,000 hours of English audio a month (600,000 minutes). Baseline scenario: pre-recorded/batch, no diarization, no custom vocabulary, no summarization. Feature-heavy scenario: the same audio with speaker diarization enabled, and, for streaming use cases, real-time transcription instead of batch. Volume discounts and committed tiers (Deepgram Growth, AssemblyAI or OpenAI enterprise agreements) are not applied; all figures are pay-as-you-go list rates.

Cost at the normalized workload

Formula: cost = minutes × per-minute rate, with add-ons summed separately where billed per minute.

Provider / modelBatch, baselineBatch + diarizationStreaming
AssemblyAI Universal-2$1,500Diarization reported bundled; no confirmed added per-minute rate—
OpenAI gpt-4o-mini-transcribe (DEPRECATED Aug 26, 2026; available until Feb 26, 2027)$1,800No native diarization on this model—
AssemblyAI Universal-3.5 Pro$2,100Diarization reported bundled$4,500 (Pro Realtime)
Deepgram Nova-3$2,580$5,220 ($2,580 + $2,640 diarization add-on)$4,620 (base); $7,260 with diarization add-on
OpenAI gpt-transcribe (current recommended default)$2,700Confirm diarization support directly—
OpenAI gpt-4o-transcribe (DEPRECATED Aug 26, 2026; available until Feb 26, 2027)$3,600No native diarization on this model—
OpenAI gpt-live-transcribe (current realtime model)——$10,200

AssemblyAI's Universal-2 is the cheapest baseline batch option modeled here, followed by Deepgram at $1,080 more. OpenAI's current recommended gpt-transcribe, at $2,700, sits between Deepgram and the deprecated gpt-4o-transcribe's $3,600; the deprecated gpt-4o-mini-transcribe's $1,800 remains a cheaper transitional figure for teams still migrating off it, but is not the number to budget against for a new integration. Deepgram's per-minute diarization add-on is the single largest feature surcharge found in this comparison, roughly doubling its batch bill. For streaming, OpenAI's current gpt-live-transcribe is more than twice Deepgram's or AssemblyAI's streaming rate.

Why "feature-heavy" moves Deepgram the most

Deepgram publishes diarization, redaction, keyterm prompting and entity detection as separate per-minute add-ons on top of the base transcription rate, each with its own price; stacking several (diarization plus redaction plus keyterm prompting, for instance) has been reported to more than double the effective per-minute cost relative to bare Nova-3. AssemblyAI's and OpenAI's core batch models are reported to bundle diarization (where offered) into the base rate rather than metering it separately, which is why the "batch + diarization" column shows no additional charge for them — confirm current behavior directly, since add-on structures change.

Sensitivity

  1. Diarization and add-on features. The single largest lever for Deepgram specifically; confirm whether the add-on applies on your committed tier, since Growth pricing reportedly reduces or bundles it differently than pay-as-you-go.
  2. Batch versus streaming. Every vendor charges roughly 1.5 to 3 times its batch rate for real-time transcription; OpenAI's gap is the widest.
  3. Which model tier, and whether it is on OpenAI's deprecation path. AssemblyAI's Universal-2 versus Universal-3.5 Pro is a 40% price difference for a reported accuracy gain; on OpenAI, the deprecated gpt-4o-mini-transcribe versus gpt-4o-transcribe was a 2x difference, now superseded by the single current gpt-transcribe rate of $0.0045/minute.
  4. Committed versus pay-as-you-go pricing. Deepgram's Growth tier is reported to offer up to 20% off; similar volume discounts from AssemblyAI and OpenAI are negotiated, not published.

Budgeting traps

  • Assuming whisper-1 or gpt-4o-transcribe/gpt-4o-mini-transcribe are OpenAI's current recommended models. All four (whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize) were deprecated August 26, 2026, with shutdown scheduled for February 26, 2027. The current recommended path is gpt-transcribe for batch and gpt-live-transcribe for realtime.
  • Pricing Deepgram from the base Nova-3 rate alone. Diarization and other add-ons are separately metered and can roughly double the real bill.
  • Comparing a feature-complete AssemblyAI or OpenAI quote to a bare Deepgram rate. Whether diarization is bundled or separately billed differs by vendor and should be confirmed for the specific model in use.
  • Using a streaming rate to estimate a batch job, or vice versa. The gap is 50% to 180% across these three vendors.
  • Ignoring committed-tier discounts at genuine 10,000-hour-a-month volume. Deepgram's Growth tier and vendor-negotiated agreements can move the real price meaningfully below pay-as-you-go.

What to ask before you buy

Confirm each vendor's current model name and rate directly, since all three have moved pricing and model generations more than once in 2026, and OpenAI specifically has a full deprecation and shutdown schedule in effect for its prior transcription model family. Ask specifically whether diarization, redaction and other features you need are bundled into the base rate or billed as separate per-minute add-ons, and get a committed-tier quote at your real volume before budgeting off pay-as-you-go rates.