Qwen 2026: What Does Alibaba's AI Actually Cost?
Qwen has a free consumer interface and a separate, metered API. Qwen Studio describes its chat assistant as free to use. Building a product with Alibaba Cloud Model Studio is different: input and output tokens are charged separately, and the same model can have a different list price depending on the region and service deployment scope. A free chat session is not an API quota.
For a concrete starting point, the Singapore/International standard list rate for qwen3.8-flash is $0.15 input / $0.47 output per million tokens. qwen3.8-max is $2.00 / $6.00 there. The corresponding Global-scope rates displayed for several other regions are $0.113 / $0.382 and $1.65 / $4.951. Those are not interchangeable checkouts: region affects data location, keys, endpoint configuration, model availability and possibly compliance requirements.
Chat, API and your own hardware
| Access path | What it costs | What to check |
|---|---|---|
| Qwen Studio | Advertised as free for personal chat | The public page does not promise a particular daily message allowance or production service level. |
| Model Studio API | Pay as you go by model, token type and region | Use the exact regional table and service scope for the endpoint you will call. |
| Open-weight Qwen deployed by you | No Alibaba per-token API invoice for your own inference | Hardware, storage, operations and the license of the specific released checkpoint still matter. |
Avoid treating the brand name as a license. Alibaba publishes both hosted proprietary models and open-weight releases; a self-hosted Qwen model is not necessarily the same checkpoint or quality level as qwen3.8-max on the API.
Current Model Studio text rates
The table below uses Alibaba's official standard list prices, in USD per one million uncached input or output tokens. Singapore/International is a different deployment scope from Global in Hong Kong, Frankfurt, Virginia or Tokyo. The Max and Flash rows shown do not have a prompt-size price step within their listed 1M-token range; Plus does.
| Model and scope | Input | Output | Per-request prompt tier |
|---|---|---|---|
qwen3.8-flash, Singapore/International | $0.15 | $0.47 | Up to 1M input tokens |
qwen3.8-flash, listed Global scopes | $0.113 | $0.382 | Up to 1M input tokens |
qwen3.7-plus, Singapore/International | $0.40 | $1.60 | Up to 256K input tokens; standard list price before any console promotion |
qwen3.7-plus, Singapore/International | $1.20 | $4.80 | Above 256K and up to 1M input tokens |
qwen3.8-max, Singapore/International | $2.00 | $6.00 | Up to 1M input tokens |
qwen3.8-max, listed Global scopes | $1.65 | $4.951 | Up to 1M input tokens |
Alibaba notes that its documentation shows standard prices and that promotions belong in the console. Do not bake a limited-time discount into a recurring production budget. For tiered models such as Plus, Alibaba bills all tokens in a request at the tier reached by that request's input size, not just the tokens above the threshold. Output produced in thinking mode also consumes billed output tokens.
The regional difference is real: one million input plus one million output tokens on qwen3.8-max, spread across eligible requests, totals $8.00 on Singapore/International and $6.601 on a listed Global scope, before caching or other services. That is about 17.5% less in this example. It is not an invitation to change only a URL in production. Alibaba's region guide says each region has its own API key, endpoint and model list; selected region determines request storage, while scope determines where inference may run. Verify the data boundary and model availability first.
Cache and free-quota details that change the bill
Model Studio's free-quota policy grants eligible new users a per-model trial quota only in Singapore with an International deployment scope. The pricing table lists one million tokens for the models above, valid for 90 days under its stated start-date rules. Other regions and scopes do not inherit that free quota. Alibaba says pay-as-you-go charges begin after the quota is used unless the account enables its Free Quota Only control.
Context caching can reduce repeated input costs, but only matching, eligible input is billed as a hit. Explicit cache creation and hits have their own rules; importantly, the documentation identifies exceptions to its typical percentage discounts for some Qwen3.8 models. Do not assume every cache hit is exactly 10% of the ordinary input rate. The rate card above deliberately excludes cache discounts. Batch discounts also apply only where a model's row explicitly supports batch processing and cannot be stacked with context-cache discounts.
Two budgets, same work
For the examples below, use ordinary real-time requests, no cached tokens, and no console promotion. Each request remains below the relevant context threshold unless noted.
Short document answer: 10,000 input tokens and 2,000 output tokens.
| Model/scope | Calculation | Cost per answer |
|---|---|---|
qwen3.8-flash, Singapore/International | 0.01 x $0.15 + 0.002 x $0.47 | $0.00244 |
qwen3.8-flash, listed Global scope | 0.01 x $0.113 + 0.002 x $0.382 | $0.001894 |
qwen3.7-plus, Singapore/International | 0.01 x $0.40 + 0.002 x $1.60 | $0.00720 |
qwen3.8-max, Singapore/International | 0.01 x $2.00 + 0.002 x $6.00 | $0.03200 |
Five hundred long-document summaries: 100,000 input tokens and 5,000 output tokens per document, one request per document. At Singapore/International list prices, qwen3.7-plus costs $0.048 per document or $24 for 500; qwen3.8-max costs $0.23 per document or $115 for 500. Plus is in its lower tier because each request has 100K input tokens. If a single Plus request instead carries 300K input and 20K output tokens, it crosses the 256K threshold and costs $0.456 at $1.20 input / $4.80 output, not $0.152 at the lower tier.
These comparisons price tokens, not answer quality. The cheapest checkpoint that meets your accuracy target is the useful cheapest option. Tool calls, search, images and output verbosity can change a real invoice.
The practical choice
Use Qwen Studio if you need chat and its free experience suits your usage. For an app, choose the specific model and regional scope before comparing prices. Start with Flash for repeatable low-cost tasks, test Plus when Flash misses quality requirements, and pay for Max only where that evaluation earns its higher rate. If you need a particular data location, a lower Global price is not a substitute for that requirement. Compare equivalent short workloads in Grok vs Qwen vs DeepSeek vs ChatGPT and the separate DeepSeek API cost guide.
For a wider view of how Chinese providers compare on price, cache behavior and region, see the Chinese AI API price comparison; its StepFun pricing analysis gives the provider-specific detail.
Frequently asked questions
Is Qwen free? The Qwen Studio chat page advertises free use. Model Studio API usage is separately metered after any eligible trial quota.
Does region still change actual cost? Yes. Alibaba's current regional rate card lists different rates for the same Qwen3.8 model under Singapore/International and listed Global scopes. Region and scope also change where requests are stored or processed.
Can a free API quota make a non-Singapore endpoint free? No. Alibaba restricts this new-user quota to eligible Singapore/International models.