Qwen Pricing Explained (2026): What Alibaba's AI Actually Costs
Qwen has one of the most confusing pricing stories in AI, not because Alibaba hides anything, but because "Qwen" is really three different products wearing one name. There's the free consumer chat app. There's the developer API, billed by the token through Alibaba Cloud Model Studio, which used to have a genuinely free tier and doesn't anymore. And there's the open-weight model family, which you can download and run yourself for the cost of your own hardware.
Most "is Qwen free?" articles answer only the first question and stop. That's how you end up budgeting a product roadmap around a free API tier that Alibaba shut down back in April 2026, or missing the fact that the flagship model's price has already changed twice since midyear. This guide separates the three products cleanly, prices each one using Alibaba's own current rate card, and answers the question that actually matters for a buyer: is Qwen one of the cheapest serious AI options in 2026, and for whom?
Quick answer: Qwen Chat, the consumer app, is free and has no subscription tier — a genuine rarity next to ChatGPT Plus or Claude Pro. The developer API is not free; it's priced per token, with the current flagship (Qwen3.8-Max) at $2.00 per million input tokens and $6.00 per million output tokens on Alibaba's international endpoint — cheap by frontier standards, but not zero. The free developer tier that many older guides still describe closed on April 15, 2026. What replaced it is a 90-day, 1-million-token-per-model trial, which is generous for evaluation but not a permanent free ride.
The three things people mean when they say "Qwen"
Before any numbers, it's worth being precise about what you're actually pricing, because the three paths have almost nothing in common cost-wise.
- Qwen Chat (chat.qwen.ai, plus iOS, Android and macOS apps) — a free consumer assistant, comparable in spirit to the ChatGPT or Gemini apps, but with no paid subscription tier sitting above it.
- Alibaba Cloud Model Studio (formerly branded DashScope) — the developer API. This is where Qwen actually makes money from most non-Chinese users, billed per input and output token, with pricing that varies by region, model and even time of day.
- Open-weight Qwen3 models — most of the Qwen3 family (from the 0.6B models up through the 235B-A22B mixture-of-experts model) is released under the Apache 2.0 license on Hugging Face. You can run the exact model the API serves, on your own GPUs, for $0 in licensing fees — you just pay for the compute.
A fourth option has appeared during 2026 for people who don't want to think in tokens at all: Alibaba's Token Plan, a prepaid subscription that converts a flat monthly fee into a credit balance spendable across the whole Model Studio catalog (text, image, video and audio models together). More on that below.
Is the Qwen Chat app actually free?
Yes, and unusually cleanly so. Qwen Chat requires no sign-up to try, carries no advertised rate limits, and — as of September 2026 — has no paid consumer tier at all. There is no "Qwen Plus" subscription sitting next to it the way ChatGPT Plus sits next to the free ChatGPT tier. If you only ever use Qwen through the chat app or its mobile clients, your bill is $0, indefinitely, with no credit card on file and no trial clock running out.
That's a genuinely different shape from every Western competitor. ChatGPT, Claude and Gemini all use a "free tier gates you toward a $20/month subscription" model. Qwen doesn't have that upsell mechanism in the consumer product at all — Alibaba's monetization for individuals happens entirely on the API and infrastructure side, not the chat interface. The trade-off is what you'd expect: the free chat app doesn't expose fine-grained model selection or the highest-context settings the API offers, and heavy or automated use will eventually run into fair-use throttling like any free web service. But there is no dollar figure anywhere in that experience, and there doesn't appear to be one coming.
The API: what actually costs money
This is where "Qwen is basically free" stops being true. The developer API — the thing you'd use to build a product, a chatbot backend, or an automated workflow — is billed per token through Alibaba Cloud Model Studio, and it's had a genuinely disruptive year.
The free OAuth developer tier was discontinued on April 15, 2026. For over a year before that, developers could get a limited number of free requests per day (reported at 1,000, later cut to 100) without paying anything. That's gone. So was the Qwen Code CLI's free coding tier (2,000 requests/day), pulled around the same time. If you're reading a tutorial or blog post that promises standing free API access to Qwen, it's describing a product that no longer exists.
What replaced it is a one-time onboarding trial: new Alibaba Cloud accounts get 1 million free tokens per eligible model, valid for 90 days from account activation (or from a given model's release date, whichever is later). Because the quota is per-model rather than shared across the whole catalog, and because Alibaba runs dozens of Qwen model variants, this adds up to tens of millions of free tokens if you're willing to spread evaluation across the lineup — genuinely one of the more generous trial structures among major providers. The catch: it only applies on the Singapore (International) region endpoint. Model Studio deployments in Beijing, the US, the EU, and the "Global" routing tier carry no free quota at all.
Current Qwen API pricing, straight from Alibaba's rate card
All figures below are Alibaba Cloud Model Studio's own published standard prices (excluding time-limited promotions), verified against the official pricing page on September 14, 2026. Prices are per 1 million tokens, in USD, on the Singapore/International endpoint unless noted. "Thinking mode" pricing applies when a model's extended chain-of-thought is switched on; output is billed as chain-of-thought tokens plus the final answer combined.
| Model | Input (per 1M) | Output (per 1M) | Context | Free quota (Singapore only) |
|---|---|---|---|---|
| Qwen3.8-Max (current flagship, GA Aug 3, 2026) | $2.00 | $6.00 | 1M | 1M tokens / 90 days |
| Qwen3.7-Max (prior flagship) | $2.50 | $7.50 | 1M | 1M tokens / 90 days |
| Qwen3-Max (0–32K tokens/request) | $1.20 | $6.00 | up to 256K, tiered | 1M tokens / 90 days |
| Qwen3-Max (32K–128K tokens/request) | $2.40 | $12.00 | tiered | — |
| Qwen3-Max (128K–256K tokens/request) | $3.00 | $15.00 | tiered | — |
| Qwen-Max (legacy, no tiering) | $1.60 | $6.40 | — | 1M tokens / 90 days |
| Qwen3.7-Plus (0–256K tokens, promo) | $0.40 | $1.60 | 1M, tiered | 1M tokens / 90 days |
| Qwen3.7-Plus (256K–1M tokens, promo) | $1.20 | $4.80 | tiered | — |
| Qwen-Plus (0–256K tokens, non-thinking / thinking) | $0.40 | $1.20 / $4.00 | 1M, tiered | 1M tokens / 90 days |
| Qwen3.8-Flash | $0.15 | $0.47 | 1M | 1M tokens / 90 days |
CALCULATION NOTE: Qwen3.7-Plus's $0.40/$1.60 row is a limited-time 20% discount off Alibaba's list price of $0.50/$2.00 — both figures come directly from the official pricing page, which explicitly labels it "Limited-time 20% off." Budget off the list price if you want a stable multi-year estimate; the promotional rate is what you'll actually be billed today.
Regional pricing is not the same price
This is the detail almost every third-party "Qwen pricing" roundup gets wrong by quoting only one region. Alibaba runs genuinely different rate cards in China (Beijing), Singapore (International), Hong Kong, Germany (Frankfurt/EU), the US (Virginia) and Japan (Tokyo), and the differences aren't rounding errors:
| Region | Qwen3.8-Max input/output (per 1M) | Free quota |
|---|---|---|
| Singapore (International) | $2.00 / $6.00 | Yes — 1M tokens/model |
| China (Beijing) | $1.65 / $4.951 | No |
| Hong Kong, Germany, US, Japan ("Global" scope) | $1.65 / $4.951 | No |
Counterintuitively, the region with the free trial is the most expensive per-token region once you're paying. If you're a non-Chinese developer choosing a Model Studio region purely to minimize your steady-state bill, the "Global" deployment scope (served from Hong Kong/Frankfurt/Virginia/Tokyo infrastructure) is roughly 17–18% cheaper per token than Singapore — you just give up the free evaluation quota to get there. For a team that will genuinely spend real money every month, running the free trial in Singapore first and then migrating production traffic to a Global-scope endpoint is a defensible strategy, not a loophole.
Cheaper ways to pay for the same tokens
Three mechanisms can materially cut the sticker price above, and all three come from Alibaba's own documentation:
- Batch inference. For models that support asynchronous batch calls (used for non-real-time work like bulk classification, dataset labeling or overnight report generation), both input and output are billed at 50% of the real-time rate. On Qwen3-Max, for example, batch input drops from $1.20 to $0.60 per million tokens and batch output from $6.00 to $3.00.
- Context caching. Cache-hit input tokens are billed at a steep discount versus the standard input rate — cache creation runs at 125% of the standard rate, but a cache hit is typically billed around 10% of it. For a chatbot or coding assistant that resends a long, largely-static system prompt or codebase context on every turn, this is often the single biggest lever available, because it applies automatically once the underlying content repeats.
- Batch and caching don't stack. Alibaba's documentation is explicit that these two discounts apply to different parts of a request and are not combined on the same tokens — pick whichever fits your workload's actual bottleneck (throughput vs. repeated context).
The subscription options: Token Plan and Coding Plan
Since mid-2026, Alibaba has offered two flat-fee alternatives to pure pay-as-you-go, aimed at people who don't want to manage a running token meter:
- Model Studio Token Plan (Personal Edition). A prepaid monthly subscription that converts your fee into spendable credits across the entire Model Studio catalog — text, image, video and audio models together — at up to roughly 3x the usage you'd get from spending the same amount at pay-as-you-go rates, according to Alibaba's own launch announcement. Early-bird pricing started at $6/month. This is the cleanest option for an individual power user who wants one predictable bill instead of a variable one, though the "up to 3x value" framing is Alibaba's own marketing number rather than an independently audited multiplier, so treat it as an upper bound, not a guarantee for your specific mix of models.
- Coding Plan. A fixed monthly subscription for the Qwen Code CLI, metered by number of model invocations rather than raw tokens — the direct replacement for the free coding tier that was discontinued in April 2026. Alibaba does not publish a single simple headline price for this plan on its public pricing page as of this writing; developers configuring Qwen Code are shown the current rate inside the console at sign-up. If a fixed invocation count matters more to your budgeting than a token count, it's worth checking the live console price before committing, since — unlike the token rate card — this number isn't locked into Alibaba's published documentation.
Self-hosting: the real cost floor
Because the core Qwen3 model family — from the small 0.6B models up through the 235B-A22B mixture-of-experts flagship — ships under the Apache 2.0 license on Hugging Face, you can legally download and run the exact model Alibaba serves through its API, on your own infrastructure, for $0 in licensing cost. What you pay instead is GPU rental or purchase, storage, networking, and the engineering time to deploy and maintain the stack — none of which is trivial. A dense 70B-class model needs roughly 80GB+ of VRAM for full-precision inference; the mixture-of-experts models are more forgiving because they only activate a fraction of their total parameters per token, but they still require real hardware. Self-hosting only beats the API on cost at genuinely high, sustained volume — the crossover point where a rented or owned GPU cluster pays for itself against Qwen-Plus's per-token price is usually measured in hundreds of millions of tokens a month, not thousands. Below that scale, the API is cheaper once you account for your own engineering time.
One licensing nuance worth flagging honestly: not every future Qwen release is guaranteed to ship under Apache 2.0. At least one preview-stage model has been reported carrying a more restrictive "Qwen Community License" instead, which imposes commercial-hosting conditions that deserve legal review before you build a redistribution or managed-service plan around it. Check the specific model card before assuming blanket Apache 2.0 terms across the whole future Qwen catalog.
Real cost scenarios: what you'd actually pay
These are SCENARIO / PLANNING ASSUMPTIONS — illustrative workloads, not promises from Alibaba about what "the average user" spends. Use them to sanity-check your own numbers, not as a guarantee.
Scenario 1 — Casual consumer
You use Qwen Chat in the browser or the phone app for everyday questions, writing help and the occasional image description. Cost: $0/month, indefinitely. There is no meter running and no tier to accidentally upgrade into.
Scenario 2 — Solo developer prototyping an app
You're testing a customer-support summarization tool: roughly 15 million input tokens and 4.5 million output tokens a month, split across a couple of Qwen models while you evaluate.
- On the 90-day free trial (Singapore region, per-model quota): effectively $0, as long as your total usage per model stays under 1 million tokens during the trial window — easy to do while prototyping, harder once you're iterating heavily on one model.
- On Qwen3.7-Plus after the trial (promo rate, ≤256K context): 15M × $0.40 + 4.5M × $1.60 = $6.00 + $7.20 = $13.20/month.
- On Qwen3.8-Max for the same volume: 15M × $2.00 + 4.5M × $6.00 = $30.00 + $27.00 = $57.00/month.
The gap between those two rows is the single most important number in this article: choosing the mid-tier Plus model over the flagship Max model, for a workload that doesn't need frontier reasoning, is a roughly 4.3x cost difference for the same token volume.
Scenario 3 — Production coding agent, heavy usage
A small team runs an internal coding assistant that processes roughly 200 million input tokens and 50 million output tokens a month on Qwen3.8-Max, with a large, mostly-repeated codebase context that hits the cache on about 70% of input tokens.
- Without caching: 200M × $2.00 + 50M × $6.00 = $400 + $300 = $700/month.
- With ~70% cache-hit input (at roughly 10% of the standard input rate, per Alibaba's documented caching discount — approximately $0.20–$0.25 per million tokens on this model depending on the exact cache mechanics): 140M cached × ~$0.25 + 60M uncached × $2.00 + 50M output × $6.00 = $35 + $120 + $300 = ~$455/month.
That's roughly a 35% reduction from context caching alone, on a workload that's a realistic shape for a coding agent (a large, largely static system prompt and codebase snapshot resent on every call). This is a CALCULATION, not a quoted Alibaba figure — the exact cached-input rate for Qwen3.8-Max specifically should be confirmed in the console for your account, since cache pricing details can vary slightly by model and are documented separately from the headline rate card.
Qwen vs. paying for ChatGPT, Claude or Gemini
This is the comparison most readers actually want, so it's worth being precise about what's being compared. Qwen has no consumer subscription, so there's no direct "$20/month vs. $20/month" comparison to make against ChatGPT Plus, Claude Pro or Google's AI Pro tier — those bundle a chat interface, higher usage limits and API-adjacent perks into one monthly fee, and Qwen simply doesn't sell that bundle. The honest comparison is either "free chat app vs. free chat app" (where Qwen Chat holds up fine for everyday use, with the caveat that it doesn't expose the highest-context or most advanced reasoning settings the way a paid competitor tier might) or "API vs. API," where the gap is large and consistent.
| Model | Input $/1M | Output $/1M |
|---|---|---|
| Qwen3.8-Max (Singapore) | $2.00 | $6.00 |
| Qwen-Plus (promo) | $0.40 | $1.20–$1.60 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Claude Opus 5 | $5.00 | $25.00 |
At the flagship tier, Qwen3.8-Max and Claude Sonnet 5 charge an identical $2.00 per million input tokens — the difference is entirely on the output side, where Qwen is 40% cheaper. Against Claude Opus 5, Qwen3.8-Max is 2.5x cheaper on input and roughly 4x cheaper on output. Against the mid-tier Qwen-Plus, the gap against any Western flagship widens further, though that comparison isn't entirely fair: Plus is explicitly the lower-capability tier in Alibaba's own lineup, not a flagship. The realistic takeaway is that Qwen's API pricing is genuinely aggressive at the flagship level and dramatically aggressive at the mid-tier level, but you're comparing raw token cost, not proven output quality on your specific task — a cheaper token that requires more retries to get a usable answer isn't actually cheaper in practice.
Hidden costs and things that make "cheap" misleading
- The free API tier you might be planning around no longer exists. This is the single most common mistake in older content still circulating about Qwen: assuming a free developer tier that Alibaba closed on April 15, 2026.
- The free quota is region-locked. If your infrastructure sits in the US, EU, or China regions, there is no free quota at all — only the Singapore/International endpoint carries it.
- Verbosity offsets some of the savings. Independent trackers have noted that Qwen models can generate more output tokens per task than some Western competitors for equivalent instructions, which quietly eats into the headline per-token discount on output-heavy workloads. Always benchmark actual token counts on your own prompts rather than assuming list-price ratios translate directly into bill-size ratios.
- Tiered pricing by prompt length is easy to miss. Several models (Qwen3-Max, the legacy Qwen-Plus lineup) charge a materially higher rate once a single request's input crosses 32K, 128K or 256K tokens — and the entire request is billed at the higher tier's rate, not just the tokens above the threshold. A workload that occasionally sends long documents can land in a much more expensive bracket than its "typical" request would suggest.
- The Coding Plan's exact price isn't published. Unlike the token rate card, Alibaba doesn't post one headline number for the fixed-fee Coding Plan; you see it in-console. Don't budget off a number you saw in a third-party blog post for this specific plan without checking your own console.
When Qwen is genuinely the cheapest option — and when it isn't
Qwen is genuinely cheap when: you're running high-volume, output-light tasks (classification, extraction, structured data pulls) on Qwen-Flash or Qwen-Plus, where the per-token rate undercuts almost every Western equivalent by a wide margin; when your workload benefits from context caching (a large repeated system prompt or codebase); when you're comfortable evaluating for free within the 90-day, region-locked trial before committing spend; or when you're an individual who just wants a capable free chat assistant with zero subscription pressure.
The "cheap" framing gets misleading when: you need the absolute frontier tier (Qwen3.8-Max at long context, 128K–256K token prompts), where tiered pricing pushes the effective rate up toward Western flagship territory; when you're deploying outside Singapore and lose the free trial entirely; when your task needs few retries and high reliability on the first try, since a wrong or verbose answer on a "cheap" model can cost more in wasted output tokens and engineering time than a pricier model that gets it right the first time; or when you're comparing Qwen's uncapped consumer chat app against a paid competitor tier that includes genuinely different capabilities (larger context, more advanced agentic tool use) that the free tier doesn't match.
Who benefits most from Qwen's pricing
- Individual users who want a free daily assistant with no subscription anxiety — the chat app genuinely has no tier to accidentally upgrade into.
- Startups and developers building high-volume, well-defined tasks (support ticket triage, content classification, translation, first-draft generation) where Qwen-Flash or Qwen-Plus's rock-bottom per-token price directly reduces unit economics.
- Teams with infrastructure or compliance reasons to self-host, who can take the open-weight Qwen3 models under Apache 2.0 and run them without licensing fees, provided they have the GPU capacity and engineering time to justify it.
- Multilingual products, especially those serving Chinese, Japanese or Korean users, where Qwen's training data gives it a genuine quality edge that Western models often can't match at any price.
Is Qwen actually one of the cheapest serious AI ecosystems in 2026?
Yes, with real qualifications. At the mid-tier (Qwen-Plus, Qwen-Flash) it's genuinely among the cheapest API options from any major lab, Chinese or Western. At the flagship tier (Qwen3.8-Max) it's competitively priced against Claude and roughly matched or cheaper than most comparable Western frontier models, but it's no longer the free-lunch story that circulated when the OAuth free tier still existed. The consumer chat app remains a real outlier — a capable, genuinely free assistant with no subscription tier at all — and that alone makes Qwen worth a serious look for anyone who wants frontier-adjacent AI without a recurring bill. For developers, the honest framing is: cheap, not free, region-dependent, and materially cheaper than the Western majors at the tiers most projects actually need.
FAQ
Is Qwen completely free to use? The consumer chat app (chat.qwen.ai and its mobile/desktop apps) is free with no subscription tier. The developer API is not free — it's billed per token, though new accounts get a 90-day, 1-million-token-per-model trial on the Singapore region only.
When did Qwen's free API tier end? The free OAuth developer tier (previously 1,000, later 100 free requests/day) was discontinued on April 15, 2026. The Qwen Code CLI's free coding allowance ended around the same time.
How much does Qwen3.8-Max cost? $2.00 per million input tokens and $6.00 per million output tokens on Alibaba Cloud Model Studio's Singapore/International endpoint, as of Alibaba's rate card checked September 14, 2026. Pricing is roughly 17–18% lower ($1.65/$4.951) on the "Global" deployment scope covering Hong Kong, Germany, the US and Japan, which does not carry a free trial quota.
Is the Qwen API cheaper than Claude or ChatGPT? At the flagship level, Qwen3.8-Max matches Claude Sonnet 5 on input price ($2.00/1M) and is roughly 40% cheaper on output. Against higher-end models like Claude Opus 5, Qwen is several times cheaper per token. At Qwen's mid-tier (Qwen-Plus), the gap against any Western flagship is substantially wider.
Can I run Qwen for free by self-hosting? Most of the Qwen3 model family is released under the Apache 2.0 license, so there's no licensing fee to self-host. You still pay for GPU infrastructure, storage and engineering time, which only becomes cheaper than the API at genuinely high, sustained usage volumes.
What is the Qwen Token Plan? A prepaid monthly subscription (early-bird pricing from $6/month) that converts a flat fee into spendable credits across Alibaba's whole Model Studio catalog — text, image, video and audio models together — at up to roughly 3x the usage of paying the equivalent amount at standard pay-as-you-go rates, according to Alibaba's own announcement.
Does Qwen pricing differ by region? Yes, materially. The Singapore/International endpoint is the only one carrying a free trial quota, but it's also the most expensive per-token region for Qwen3.8-Max and Qwen3.7-Max once you're past the trial. Beijing, Hong Kong, Frankfurt, Virginia and Tokyo ("Global" scope) run a lower, harmonized rate with no free quota.
What happened to Qwen 2.5 pricing? Qwen 2.5 is a prior model generation; current Alibaba pricing pages are built around the Qwen3.x family (Qwen3.8-Max, Qwen3.7-Plus, Qwen3.8-Flash, etc.). If you're budgeting from a guide that still quotes Qwen 2.5 numbers, it's describing a model line Alibaba's live rate card no longer centers on.
Is there a discount for high-volume or batch workloads? Yes — asynchronous batch inference is billed at 50% of the real-time rate on supported models, and context caching can cut input costs on repeated content by roughly 90% on a cache hit. These two discounts don't stack on the same tokens.
Do I need a China-based account to use Qwen? No. International developers typically use the Singapore (International) region of Alibaba Cloud Model Studio, which is the default region carrying the free evaluation quota for non-Chinese accounts.
Schema recommendations
- Article schema: headline "Qwen Pricing Explained (2026): What Alibaba's AI Actually Costs"; datePublished/dateModified set to the September 2026 publication date; author "NotCheapAI Research Desk"; publisher NotCheapAI with logo; mainEntityOfPage set to the canonical slug.
- BreadcrumbList schema: Home → Guides → Qwen Pricing Explained (2026), matching the site's existing
/guides/hierarchy. - FAQPage schema: appropriate here — the FAQ section above is genuinely visible, user-facing Q&A content matching the on-page questions verbatim, which meets the bar for FAQPage markup.