The rate card
What a million tokens costs
List prices in US dollars, read from Alibaba Cloud's model inference pricing page. Every figure is the Singapore region under the international deployment scope, for a single request up to one million input tokens.
| Model | Request size | Input | Output |
|---|---|---|---|
| qwen3.8-max checked 2026-10-06 |
up to 1M tokens | $2.00 | $6.00 |
| qwen3.8-flash checked 2026-10-06 |
up to 1M tokens | $0.15 | $0.47 |
| qwen3.7-max checked 2026-10-06 |
up to 1M tokens | $2.50 | $7.50 |
| qwen3.7-flash checked 2026-10-06 |
up to 32K tokens | $0.030 | $0.130 |
| qwen3.7-flash same model, longer request |
32K to 256K tokens | $0.100 | $0.400 |
| qwen3.7-flash same model, longer request |
256K to 1M tokens | $0.200 | $0.800 |
Rates read and checked on October 6, 2026. These are list prices: the free quota, batch discounts and cache rates are separate, and each is covered below. Alibaba lists a single tier for both 3.8 models, covering requests up to a million input tokens. The price list says nothing about longer requests, so neither do we.
The same model is cheaper in other regions on Alibaba's own page. qwen3.8-max reads $1.65 in and $4.951 out in the Beijing table, and the tables are split by region and by deployment scope on purpose. We keep one region on this page, because two lists in one table is how a price page stops being trustworthy. If you are billed somewhere other than Singapore, open the source and read that column.
Rates are half of the question. What a month of them costs is the other half, and Qwen API pricing by workload multiplies these numbers out over four monthly volumes.
Free quota
The free quota, briefly
A new Model Studio account starts with free tokens rather than a bill, and the number people repeat is a million per model. The offer only holds while three things are true at once: the Singapore region with the international deployment scope, typically 1,000,000 tokens per model, and 90 days counted from whichever came last, activation, the model's release, or approval of your request. A dated snapshot counts as a model of its own.
The small print, what the quota does not pay for, and the switch that stops the service with a 403, error code AllocationQuota.FreeTierOnly, instead of billing you, are on Is Qwen free.
Rate limits
The ceiling on how fast you can send
What decides how many requests you get is separate from what decides the bill. On the international list the qwen3.8 models are marked Dynamic Rate Limiting rather than given fixed numbers, so Singapore has no published requests-per-minute figure for them. The Global scope rows do carry numbers: 30,000 requests per minute and 5,000,000 tokens per minute for qwen3.8-flash and qwen3.8-max, with the dated qwen3.8-max-0902 snapshot capped lower, at 150,000 tokens per minute.
Paying does not raise that ceiling either. Limits are set at the account level and are independent of billing, and a top-up only keeps the account from being suspended over an unpaid balance. If you need a higher limit, the documentation sends you to your business manager rather than to a pricing page.
Two error codes are worth keeping apart when something stops working: a 429 means a request went past the requests-per-minute or tokens-per-minute limit and usually clears within a minute, while a 403 carrying the FreeTierOnly code means the free quota is gone. Alibaba calls both rejection mechanisms rather than failures, which is the right way to read them. Whether a free call is slower than a paid one is a separate question, and it is answered on Is Qwen free.
Caching
What prompt caching does to a Qwen bill
The 3.8 rows on Alibaba's price list carry a context caching discount flag, which tells you cached input is billed below the standard input rate. The documentation does not say by how much for these particular models, and that gap is Alibaba's rather than ours. The pricing page gives an example, 125% for creating an explicit cache and 10% for hitting one, then points elsewhere for the numbers that apply; the cache page gives 20% of the input rate as the general rule, names qwen3.8-max, qwen3.8-flash and two other models as exceptions to it, and sends you to the console for their real prices.
So this page prints no cache rate for Qwen 3.8. When a number is not on a page we can link to, the cell stays empty. Batch work is clearer: the same price list marks the qwen3.8 rows with a 50% batch inference discount, meaning input and output both bill at half the listed rate for batch jobs.
Open weights
Open weights, without a per-token bill
Alibaba keeps a second list of Qwen models under the open source heading. These are weights you can download and run yourself; the prices below are what Alibaba charges to host them for you.
| Model | Request size | Input | Output |
|---|---|---|---|
| qwen3.8-2.4t-a95b checked 2026-10-06 |
up to 1M tokens | $2.00 | $6.00 |
| qwen3.8-27b checked 2026-10-06 |
up to 1M tokens | $0.50 | $3.00 |
Running one of them on hardware you already own moves the cost from the token to the machine, and our Strata guide works through that side: the memory a 125-billion-parameter Qwen model needs, and the speeds the project measured on the cards it tested. One caution, and it is the same one that guide opens with. The model Strata runs is called Qwen3.8-Flash-Next in the project's own repository, and Alibaba's price list never uses that name. We cannot show the two are the same model, so we do not put their prices in one table.
FAQ
Qwen pricing: common questions
Method
How we check these Qwen numbers
Every figure here was read out of the page it names, and every one of them links back to it. Four pages do the work: the model inference pricing page for the rates, the free-quota page for the three conditions and the small print, the rate-limit page for the ceilings and the error codes, and the context cache page for how caching is billed.
One thing to know about that price list: it changes, and it says so itself. Alibaba stamps it with a last-updated date, and on October 6, 2026 that stamp moved from October 3 to October 6 while we were working on this page. The rates above come from the later copy. That is what the date at the top of this page is for. If a rate here disagrees with the source, the source wins, and a note to support@llmomics.com gets it fixed.