The rates
The two numbers that decide your bill
Alibaba prices Qwen by the million tokens, separately for what you send and what comes back. On the Singapore list with the deployment scope International, and for requests up to a million input tokens, the two current models sit far apart.
Two habits of that list are worth knowing before you budget. Output always costs more than input, by a factor of three to four on these models, so a chatty answer is the expensive half of a request. And the price is attached to a region and a request size, not to the model name alone, which is why the same model can read a different number in Alibaba's Beijing table.
Worked examples
Four monthly bills, worked out
Each example below is a monthly volume in millions of tokens, multiplied by the rates above. Nothing here is measured from a live account; it is arithmetic on published prices, and you can check every line with a calculator.
| Workload | Tokens in / out | The arithmetic | A month |
|---|---|---|---|
| Support bot, flash | 3,000,000 in / 1,000,000 out |
3 × $0.15 = $0.45 1 × $0.47 = $0.47 |
$0.92 |
| The same bot, max | 3,000,000 in / 1,000,000 out |
3 × $2 = $6.00 1 × $6 = $6.00 |
$12.00 |
| Overnight batch, flash | 20,000,000 in / 2,000,000 out |
20 × $0.15 = $3.00 2 × $0.47 = $0.94 |
$3.94, or $1.97 as a batch |
| Coding agent, max | 20,000,000 in / 2,000,000 out |
20 × $2 = $40.00 2 × $6 = $12.00 |
$52.00 |
Column two is the monthly volume in tokens; the cost columns are that volume in millions, times the rate per million.
Read rows two and four together and the shape of a bill shows up. Same token counts, same month, and the one on max costs thirteen times as much. The rate did all of that, not the volume. Output length does something similar inside a single row: it is the expensive half of every request, so an agent that writes long answers on the cheap model can still land closer to the expensive one than its input count suggests.
One assumption holds all four rows up: that each individual request stays under the million-input-token ceiling of the tier these rates belong to. A single request above it moves to the next tier, and on the older flash model those tiers are three separate prices.
What changes the bill
Four levers, and what each is worth
Batch work is half price. Alibaba's own sentence: if a model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. That is what turns the $3.94 batch job above into $1.97, and for anything that can wait, it is the largest single discount on the list.
Caching is the other one, with a gap. Context caching discounts input tokens only, and it cannot be stacked with the batch discount. Alibaba's cache page gives 125% and 10% as the general example, then says the qwen3.8 models do not follow that rule and points at the console for their actual numbers. So the discount is real, the rate is not published, and we leave it out rather than print the general example as if it applied.
The free quota takes the first million. Every model with a free tier comes with typically 1,000,000 tokens, counted per model, usable in the Singapore region with the deployment scope International, and valid for 90 days. For a workload of the size of the examples above that is most of a month for nothing, which is why it belongs in the budget. The free-tier page has the conditions in full.
Request length sets the tier. These rates are the ones for requests up to a million input tokens. Past that, models with a longer context move into another row of the table, and on qwen3.7-flash the three tiers run $0.030 to $0.200 for input and $0.130 to $0.800 for output, per million tokens. Long-context work is priced like a different model, because that is how the list treats it.
FAQ
Qwen API pricing: common questions
Method
How these numbers were checked
The two rates and the batch sentence come from Alibaba Cloud's model inference pricing page, which is linked below and carries its own last-updated stamp; when we read it on October 6, 2026, that stamp said October 6. The cache rule comes from the Context Cache page, and the free-quota conditions from the new-users page. Every figure on this page is one of those, or arithmetic done on them.
The arithmetic is multiplication, and you are meant to redo it: monthly tokens in millions, times the rate per million. Where a total would land between cents we have stayed with volumes that do not, so nothing here is rounded. If a number here ever disagrees with Alibaba's page, their page wins, and a line to support@llmomics.com gets it fixed.