Rates and free-quota terms checked October 6, 2026

Qwen Pricing

Every rate on this page opens Alibaba Cloud's own price list, and so does every limit.

Qwen is Alibaba's model family. Model Studio is where you rent it by the token, and this page is about what that rent costs: the per-million-token rates, the free quota a new account starts with, the ceilings that come attached to it, and the one thing people usually get wrong about paying. A warning before the first number. Alibaba prices by region, and the same model can carry three different prices in three places. Everything below is Singapore, the international deployment scope, which is the list an account outside mainland China sees. Claude rates live on the home page, and running a Qwen model on your own machine is what the Strata guide is for.

What a million tokens costs

List prices in US dollars, read from Alibaba Cloud's model inference pricing page. Every figure is the Singapore region under the international deployment scope, for a single request up to one million input tokens.

Model Request size Input Output
qwen3.8-max
checked 2026-10-06
up to 1M tokens $2.00 $6.00
qwen3.8-flash
checked 2026-10-06
up to 1M tokens $0.15 $0.47
qwen3.7-max
checked 2026-10-06
up to 1M tokens $2.50 $7.50
qwen3.7-flash
checked 2026-10-06
up to 32K tokens $0.030 $0.130
qwen3.7-flash
same model, longer request
32K to 256K tokens $0.100 $0.400
qwen3.7-flash
same model, longer request
256K to 1M tokens $0.200 $0.800

Rates read and checked on October 6, 2026. These are list prices: the free quota, batch discounts and cache rates are separate, and each is covered below. Alibaba lists a single tier for both 3.8 models, covering requests up to a million input tokens. The price list says nothing about longer requests, so neither do we.

The same model is cheaper in other regions on Alibaba's own page. qwen3.8-max reads $1.65 in and $4.951 out in the Beijing table, and the tables are split by region and by deployment scope on purpose. We keep one region on this page, because two lists in one table is how a price page stops being trustworthy. If you are billed somewhere other than Singapore, open the source and read that column.

Rates are half of the question. What a month of them costs is the other half, and Qwen API pricing by workload multiplies these numbers out over four monthly volumes.

The free quota, briefly

A new Model Studio account starts with free tokens rather than a bill, and the number people repeat is a million per model. The offer only holds while three things are true at once: the Singapore region with the international deployment scope, typically 1,000,000 tokens per model, and 90 days counted from whichever came last, activation, the model's release, or approval of your request. A dated snapshot counts as a model of its own.

The small print, what the quota does not pay for, and the switch that stops the service with a 403, error code AllocationQuota.FreeTierOnly, instead of billing you, are on Is Qwen free.

The ceiling on how fast you can send

What decides how many requests you get is separate from what decides the bill. On the international list the qwen3.8 models are marked Dynamic Rate Limiting rather than given fixed numbers, so Singapore has no published requests-per-minute figure for them. The Global scope rows do carry numbers: 30,000 requests per minute and 5,000,000 tokens per minute for qwen3.8-flash and qwen3.8-max, with the dated qwen3.8-max-0902 snapshot capped lower, at 150,000 tokens per minute.

Paying does not raise that ceiling either. Limits are set at the account level and are independent of billing, and a top-up only keeps the account from being suspended over an unpaid balance. If you need a higher limit, the documentation sends you to your business manager rather than to a pricing page.

Two error codes are worth keeping apart when something stops working: a 429 means a request went past the requests-per-minute or tokens-per-minute limit and usually clears within a minute, while a 403 carrying the FreeTierOnly code means the free quota is gone. Alibaba calls both rejection mechanisms rather than failures, which is the right way to read them. Whether a free call is slower than a paid one is a separate question, and it is answered on Is Qwen free.

What prompt caching does to a Qwen bill

The 3.8 rows on Alibaba's price list carry a context caching discount flag, which tells you cached input is billed below the standard input rate. The documentation does not say by how much for these particular models, and that gap is Alibaba's rather than ours. The pricing page gives an example, 125% for creating an explicit cache and 10% for hitting one, then points elsewhere for the numbers that apply; the cache page gives 20% of the input rate as the general rule, names qwen3.8-max, qwen3.8-flash and two other models as exceptions to it, and sends you to the console for their real prices.

So this page prints no cache rate for Qwen 3.8. When a number is not on a page we can link to, the cell stays empty. Batch work is clearer: the same price list marks the qwen3.8 rows with a 50% batch inference discount, meaning input and output both bill at half the listed rate for batch jobs.

Open weights, without a per-token bill

Alibaba keeps a second list of Qwen models under the open source heading. These are weights you can download and run yourself; the prices below are what Alibaba charges to host them for you.

Model Request size Input Output
qwen3.8-2.4t-a95b
checked 2026-10-06
up to 1M tokens $2.00 $6.00
qwen3.8-27b
checked 2026-10-06
up to 1M tokens $0.50 $3.00

Running one of them on hardware you already own moves the cost from the token to the machine, and our Strata guide works through that side: the memory a 125-billion-parameter Qwen model needs, and the speeds the project measured on the cards it tested. One caution, and it is the same one that guide opens with. The model Strata runs is called Qwen3.8-Flash-Next in the project's own repository, and Alibaba's price list never uses that name. We cannot show the two are the same model, so we do not put their prices in one table.

Qwen pricing: common questions

Part of it, for a while. A new Model Studio account gets a free quota of typically 1,000,000 tokens per model, usable in the Singapore region with the service deployment scope set to International, and valid for 90 days. No other region is eligible. Once the pool is empty, calls are billed at the rates on this page, unless you have turned on Free Quota Only, which stops the service instead.
90 days, counted from whichever came last: the day you activated Model Studio, the day the model was released, or the day your request for it was approved. The countdown does not pause if you stop using the model, and any tokens left at the end expire anyway.
No. Alibaba's rate-limit page says free quota and paid calls use the same infrastructure and that the same model generates at the same speed. The quota controls one thing only: whether a request is accepted.
No. Limits are set at the Alibaba Cloud account level and are independent of billing. A top-up keeps the account from being suspended over an unpaid balance; it does not move the requests-per-minute or tokens-per-minute ceiling.
Because the price depends on the region and on the deployment scope, and Alibaba's tables are split by both. qwen3.8-max reads $2 in and $6 out on the Singapore list and $1.65 in and $4.951 out in the Beijing table. A request length tier can change the price as well, which is why qwen3.7-flash has three rates here.
That input tokens which hit the cache are billed below the standard input price. For the qwen3.8 models Alibaba does not publish the discounted rate on its documentation pages, and points to the console instead, so this page leaves the number out rather than quote the general example as if it applied.
Yes, two ways. Alibaba publishes open-source Qwen weights, and its own price list has a section for hosting them if you would rather not run them yourself. You can also run a Qwen model on a PC, which is what our Strata guide covers, as long as the memory the model needs fits the machine you have.

How we check these Qwen numbers

Every figure here was read out of the page it names, and every one of them links back to it. Four pages do the work: the model inference pricing page for the rates, the free-quota page for the three conditions and the small print, the rate-limit page for the ceilings and the error codes, and the context cache page for how caching is billed.

One thing to know about that price list: it changes, and it says so itself. Alibaba stamps it with a last-updated date, and on October 6, 2026 that stamp moved from October 3 to October 6 while we were working on this page. The rates above come from the later copy. That is what the date at the top of this page is for. If a rate here disagrees with the source, the source wins, and a note to support@llmomics.com gets it fixed.