Tencent Hunyuan 2026: API vs Open Weights — What Does It Really Cost?
Direct answer
Tencent's Hunyuan Hy4 is genuinely open-weight (Apache 2.0) and genuinely cheap on the API: at roughly $0.834/M input, $2.501/M output, $0.042/M cache-hit — the cheapest flagship-tier open model in its comparison set, cheaper than GLM-5.3, Qwen 3.8 Max, and well under a fifth of Kimi K3's output rate. But "open weight" and "cheap to self-host" are different claims: CALCULATION below, built on the model's own published parameter count, shows self-hosting Hy4 requires roughly 512GB+ of memory even at aggressive quantization — a data-center commitment, not a line item most teams can casually take on. This article models when, if ever, that commitment pays for itself against the API.
The model family
Hugging Face's model card (huggingface.co/tencent/Hy4-preview), consistent with multiple mirrors, describes Hy4 preview as 770B total parameters, 49B activated per token, an MoE architecture with 78 layers (1 dense FFN layer + 77 MoE layers, each with 256 routed experts plus 1 shared expert, top-8 routing per token), plus a separate native MTP layer (10B total / 0.7B active) for speculative decoding. License: Apache License 2.0, stated on the model card — an OSI-recognized open-source license without the revenue-share or Model-as-a-Service threshold language found in Kimi K3's bespoke license (covered elsewhere in this cluster). Released and open-sourced August 28, 2026. Its predecessor, Hy3 (295B total / 21B active, 256K context), launched as a preview on April 22–23, 2026 before a full open release lifted earlier geographic restrictions.
International availability
Tencent Cloud's international documentation site (intl.cloud.tencent.com / tencentcloud.com) describes a product line called "Tencent HY Global," with Quick Start, Purchase Guide and Billing Overview pages. The international product line covers Tencent HY 3D Global (3D asset generation) and a translation-specific text API (Hunyuan-translation); the cited materials do not list an international billing page for the general-purpose Hy3/Hy4 chat models. This suggests those models are served primarily through Tencent's mainland platform (TokenHub, cloud.tencent.com), though a dedicated international chat-model billing page may exist outside the cited materials.
API pricing
Reported rates - Two independent trackers list the same Hy4 price. Hy3's current rate is linked to Tencent's document at tencentcloud.com/document/product/1300/78937, although the cited materials do not expose readable price text:
| Model | Input | Output | Cache-hit |
|---|---|---|---|
| Hunyuan Hy4 (preview, current flagship) | $0.834/M | $2.501/M | $0.042/M |
| Hunyuan Hy3 (current, post-migration) | $0.13/M | $0.53/M | not confirmed |
| Hunyuan TurboS (older, high-speed) | $0.11/M | $0.28/M | not confirmed |
| Hunyuan A13B Instruct | $0.14/M | $0.57/M | not confirmed |
A real, documented pricing timeline for Hy3, worth knowing rather than collapsing into one number: an April 22, 2026 preview launched at $0.063/$0.21; the full Hy3 release on July 6–7, 2026 was announced at ¥1.00/¥4.00 per million (≈$0.14/$0.56 at the ¥7.2=$1 rate used here — an approximate conversion, not Tencent's own USD price); the current confirmed rate as of two independent September 2026 snapshots is $0.13/$0.53. This is also a real migration, not just a price change: Tencent's own product documentation (cloud.tencent.com/document/product/1823/130060) states directly that all calls to model ID hy3-preview inside its Token Plan are now automatically routed to the full hy3 model — the preview model ID still works, but it isn't actually serving the preview anymore.
Self-hosting economics
PROVIDER FACT (from the model's own published specifications): Hy4 is a 770B-parameter model. MODELING ASSUMPTION, using standard quantization math independent sources applied to this exact model: a Q4 GGUF quantization lands around 450GB, meaning realistically 512GB of system/GPU memory is the practical minimum to run it at all; a 256GB machine reportedly only fits 1–2 bit quantization, a level of compression most sources consider impractical for production quality. This is a data-center-class deployment, not a workstation one.
MODELING ASSUMPTION — GPU cost, explicitly using a range given the underlying uncertainty: a memory footprint of this size typically requires multiple high-memory GPUs (e.g., 8x 80GB-class accelerators, or equivalent), which on current cloud GPU rental markets runs very roughly $15–30/hour depending on provider, region, and commitment terms — we are giving a wide range here deliberately rather than a false-precision single number, because Exact cloud GPU pricing varies by provider, so the $15–30/hour range is an estimate. Owned-hardware and electricity scenarios are not included because their costs depend on hardware life, utilization and local power rates.
Break-even calculation
CALCULATION, using the API rate above and the wide self-hosting cost range as bounds:
At $0.834/$2.501 per million tokens, the API cost of Workload C (400M input + 100M output/month) is: (400×$0.834) + (100×$2.501) = $333.60 + $250.10 = $583.70/month.
Against a self-hosted GPU rental cost of, say, $20/hour run continuously (730 hours/month) = $14,600/month — the API is dramatically cheaper at this volume, by roughly 25x. Self-hosting only starts to make sense at a volume where the GPU rental is running at high utilization on a workload large enough to justify the fixed hardware commitment — which, at $14,600/month in hosting cost alone, requires a token volume in the tens of billions per month to approach API-equivalent cost per token, an extreme-scale use case far beyond typical single-product API usage. We are not computing an exact break-even token count here, because it depends entirely on the actual GPU rental rate secured (which varies by 2x or more across providers) — stating a precise number would manufacture false certainty the underlying inputs don't support.
The honest conclusion: for the overwhelming majority of real workloads, Tencent's own hosted API is cheaper than self-hosting Hy4, despite the model being free to download. Self-hosting becomes rational only at a scale, and with negotiated infrastructure pricing, that this article's available data cannot precisely bound — treat any specific "self-hosting break-even point" you see elsewhere with real skepticism unless it discloses its GPU rental assumption explicitly.
Real workload calculations (API only)
CALCULATION, Hy4 confirmed rate:
| Workload | Formula | Cost |
|---|---|---|
| A — Everyday (1M in, 250K out) | (1×$0.834) + (0.25×$2.501) | $1.459 |
| B — Coding/reasoning (10M in, 3M out) | (10×$0.834) + (3×$2.501) | $15.84 |
| C — High volume (400M in, 100M out/mo) | (400×$0.834) + (100×$2.501) | $583.70 |
BUILD & EARN scenario
MODELING ASSUMPTION: a document-analysis SaaS for Chinese-language legal or financial documents (where Hunyuan's Chinese-language grounding is the actual reason to select it) running Workload C-equivalent volume: $583.70/month in model cost. At an illustrative 50 enterprise clients paying $200/month each ($10,000/month revenue, not a verified market rate), model cost is under 6% of revenue.
FAQ
Is Hunyuan really free since it's open-weight? The weights are free to download under Apache 2.0, but running them requires hardware capable of holding a 770B-parameter model in memory — realistically 512GB+ — which is a real, ongoing infrastructure cost, not a one-time download.
At what point does self-hosting Hy4 beat the API? This article could not responsibly compute a precise number — it depends on your actual negotiated GPU rental rate, which varies by 2x or more across providers. At typical cloud GPU rates, the API remains cheaper for all but extreme-scale, sustained-high-utilization workloads.
Is Hy4 actually the cheapest flagship-tier open model? Per one detailed comparison, yes, among GLM-5.3, Qwen 3.8 Max, and Kimi K3 on output price — though DeepSeek V4 Flash and GLM-5.3-Flash (both smaller/cheaper-tier models, not flagship-for-flagship comparisons) are cheaper still.
Did Hunyuan Hy3's price really change three times in one year? Yes, per this research: an April 2026 preview at $0.063/$0.21, a July 2026 full release at roughly $0.14/$0.56, and a current rate of $0.13/$0.53 — alongside an actual model migration where the "preview" model ID now silently serves the full model instead.
Is Hy4's Apache 2.0 license confirmed, or just reported? Confirmed directly — Hugging Face's own model card for tencent/Hy4-preview states "released under the Apache License 2.0" explicitly, with no revenue-share or usage-threshold conditions like Kimi K3's license carries.
Can international developers access Hy3/Hy4 through an official non-mainland Tencent platform? We found no evidence of one. Tencent's international site hosts a separate "Tencent HY Global" line, but it only covers 3D generation and a translation-specific API — not general chat/coding models. Unlike Doubao (covered elsewhere on NotCheapAI, where BytePlus ModelArk is a confirmed official international alternative), Hunyuan's general-purpose models don't appear to have an equivalent dedicated international billing page yet.