MiniMax AI Pricing 2026: What Text, Voice and Hailuo Video Actually Cost
MiniMax publishes separate meters for text, speech and video. Its current USD API schedule lists MiniMax-M3 at $0.30 per million input tokens and $1.20 per million output tokens for contexts up to 512K, speech synthesis at $60 per million characters for Turbo or $100 per million characters for HD, and H3 video at $0.08 per second at 768p or $0.13 per second at 2K. The key budgeting difference is scale: text inference can cost cents for a modest workload, while generated video is billed by its duration and quickly dominates a combined product bill.
These are API rates, not the prices or usage limits of MiniMax's Hailuo consumer app. Keep those products separate when estimating a project. (MiniMax API pricing)
Current API rates
| Workload | Current published rate | Billing unit |
|---|---|---|
| MiniMax-M3, context up to 512K | $0.30 input / $1.20 output | Per 1M tokens |
| MiniMax-M3, context above 512K | $0.60 input / $2.40 output | Per 1M tokens |
| MiniMax-M2.7 | $0.30 input / $1.20 output | Per 1M tokens |
| MiniMax-M2.7-highspeed | $0.60 input / $2.40 output | Per 1M tokens |
| Speech 2.8 Turbo | $60 | Per 1M characters |
| Speech 2.8 HD | $100 | Per 1M characters |
| MiniMax H3 video, 768p | $0.08 | Per generated second |
| MiniMax H3 video, 2K | $0.13 | Per generated second |
MiniMax marks M3's first context tier as permanently 50% off relative to the struck-through standard rates on its pricing page. The price shown above is the payable rate currently displayed, not an additional discount to apply a second time. Long prompts above 512K use the higher tier. Confirm the model and context tier selected in your account before forecasting a large workload.
Speech details
Turbo and HD are different quality/latency options, not interchangeable character prices. The official schedule bills them per million characters. A million characters of Turbo synthesis therefore costs $60 before any separate voice, platform, or storage charges. Do not convert this to “per minute” without stating an assumed speaking rate and language; character density varies substantially.
Video details
H3 is metered by output duration. A five-second 768p clip costs $0.40; a five-second 2K clip costs $0.65. A ten-second clip costs $0.80 or $1.30 respectively. Reference media and regeneration can add charges: the same pricing page gives separate rules for reference images and for regenerating a 768p result at 2K. Include those only when your workflow actually uses them.
What sample workloads cost
The following calculations use the published USD rates and exclude taxes, retries, and optional tools.
Text assistant. For 10 million input tokens and 2 million output tokens on M3 within the first context tier:
(10 × $0.30) + (2 × $1.20) = $5.40
Speech generation. For 100,000 characters using Speech 2.8 Turbo:
100,000 ÷ 1,000,000 × $60 = $6.00
Video product. For 1,000 five-second clips at 768p:
1,000 × 5 × $0.08 = $400
The matching 2K workload is 1,000 × 5 × $0.13 = $650. The 2K setting costs $250 more for that volume before regeneration or input-media charges. A combined product should model these meters independently; an “all-in MiniMax cost” without expected text tokens, speech characters and video seconds is not meaningful.
Subscription plans are a separate purchase
MiniMax also offers token and resource-pack subscriptions through its platform. A subscription is not a flat-rate guarantee of unlimited API calls: the official page separates language-model token pricing from video and audio resource packs, and API consumption remains governed by the applicable quota and rate card. The consumer Hailuo app is another product with separate availability and subscription terms. Check the relevant account's checkout and usage panel rather than applying one product's monthly fee to another product's API bill.
Is MiniMax cheap for a production app?
For text-only use, calculate input and output separately. Repeated prompts, long-context tiers, priority processing, and tool calls can change the result. For speech, measure the characters actually sent to synthesis rather than estimating from audio minutes alone. For video, duration and resolution are the dominant variables: at the published H3 rate, moving from 768p to 2K increases a five-second clip from $0.40 to $0.65.
The practical comparison is not “one provider versus three subscriptions.” Compare equivalent model, duration, resolution, input assets and commercial terms. Hailuo consumer access, MiniMax API access, and any self-hosted model are different offerings with distinct meters and limits.
Frequently asked questions
How much does MiniMax-M3 cost?
The current API page lists $0.30 per million input tokens and $1.20 per million output tokens for context up to 512K. Above that context tier, the listed rates are $0.60 and $2.40 respectively.
How is MiniMax speech synthesis billed?
Speech 2.8 Turbo is listed at $60 per million characters and Speech 2.8 HD at $100 per million characters. The character meter is not a fixed audio-minute price.
How much does MiniMax H3 video cost?
The current USD schedule lists $0.08 per generated second at 768p and $0.13 per second at 2K. Thus a five-second output costs $0.40 or $0.65 before optional input-material or regeneration charges.
Are API prices the same as Hailuo's subscription prices?
No. Hailuo is the consumer video application; the figures in this guide are MiniMax API usage rates. Subscription plans, quotas and regional availability must be checked separately for the product being purchased.
Does one monthly subscription cover every modality?
Do not assume so. MiniMax separates token billing from audio/video resource packs and publishes different quotas and billing rules. Verify what the chosen plan includes in the account before relying on it for production usage.