Rates and worked examples checked October 6, 2026

Qwen API Pricing: What a Workload Costs

Two rates do the work. Everything below is those two numbers multiplied by a month of use.

The full rate card, including the older generation and the request-length tiers, lives on our Qwen pricing page. This page takes the two rates that cover most jobs and runs four monthly workloads through them, so you can see the bill before you write any code. The free part, which lasts 90 days in one region, is on Is Qwen free.

The two numbers that decide your bill

Alibaba prices Qwen by the million tokens, separately for what you send and what comes back. On the Singapore list with the deployment scope International, and for requests up to a million input tokens, the two current models sit far apart.

qwen3.8-max
$2 in / $6 out
Per million tokens, requests up to 1M input tokens. Alibaba's price list.
qwen3.8-flash
$0.15 in / $0.47 out
Per million tokens, same request size. Same source.
The gap between them
About 13×
On input, $2 against $0.15. Output keeps roughly the same ratio, $6 against $0.47. Model choice moves the bill more than any tuning does.

Two habits of that list are worth knowing before you budget. Output always costs more than input, by a factor of three to four on these models, so a chatty answer is the expensive half of a request. And the price is attached to a region and a request size, not to the model name alone, which is why the same model can read a different number in Alibaba's Beijing table.

Four monthly bills, worked out

Each example below is a monthly volume in millions of tokens, multiplied by the rates above. Nothing here is measured from a live account; it is arithmetic on published prices, and you can check every line with a calculator.

Workload Tokens in / out The arithmetic A month
Support bot, flash 3,000,000 in
/ 1,000,000 out
3 × $0.15 = $0.45
1 × $0.47 = $0.47
$0.92
The same bot, max 3,000,000 in
/ 1,000,000 out
3 × $2 = $6.00
1 × $6 = $6.00
$12.00
Overnight batch, flash 20,000,000 in
/ 2,000,000 out
20 × $0.15 = $3.00
2 × $0.47 = $0.94
$3.94, or $1.97 as a batch
Coding agent, max 20,000,000 in
/ 2,000,000 out
20 × $2 = $40.00
2 × $6 = $12.00
$52.00

Column two is the monthly volume in tokens; the cost columns are that volume in millions, times the rate per million.

Read rows two and four together and the shape of a bill shows up. Same token counts, same month, and the one on max costs thirteen times as much. The rate did all of that, not the volume. Output length does something similar inside a single row: it is the expensive half of every request, so an agent that writes long answers on the cheap model can still land closer to the expensive one than its input count suggests.

One assumption holds all four rows up: that each individual request stays under the million-input-token ceiling of the tier these rates belong to. A single request above it moves to the next tier, and on the older flash model those tiers are three separate prices.

Four levers, and what each is worth

Batch work is half price. Alibaba's own sentence: if a model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. That is what turns the $3.94 batch job above into $1.97, and for anything that can wait, it is the largest single discount on the list.

Caching is the other one, with a gap. Context caching discounts input tokens only, and it cannot be stacked with the batch discount. Alibaba's cache page gives 125% and 10% as the general example, then says the qwen3.8 models do not follow that rule and points at the console for their actual numbers. So the discount is real, the rate is not published, and we leave it out rather than print the general example as if it applied.

The free quota takes the first million. Every model with a free tier comes with typically 1,000,000 tokens, counted per model, usable in the Singapore region with the deployment scope International, and valid for 90 days. For a workload of the size of the examples above that is most of a month for nothing, which is why it belongs in the budget. The free-tier page has the conditions in full.

Request length sets the tier. These rates are the ones for requests up to a million input tokens. Past that, models with a longer context move into another row of the table, and on qwen3.7-flash the three tiers run $0.030 to $0.200 for input and $0.130 to $0.800 for output, per million tokens. Long-context work is priced like a different model, because that is how the list treats it.

Qwen API pricing: common questions

On the Singapore list with the deployment scope International, and for requests up to a million input tokens, qwen3.8-max reads $2 for input and $6 for output, while qwen3.8-flash reads $0.15 and $0.47. Those are per million tokens, and the two are billed separately. The full card, including the older generation and the request-length tiers, is on our Qwen pricing page.
Count your monthly input tokens and your monthly output tokens, turn each into millions, multiply by the two rates, and add the two results. A support bot sending 3,000,000 input and 1,000,000 output tokens a month on qwen3.8-flash comes to $0.45 plus $0.47, or $0.92. The same volume on qwen3.8-max comes to $12.00.
That is how the price list is built rather than a rule we can explain for Alibaba. On the current models output costs three to four times input: $6 against $2 on qwen3.8-max, and $0.47 against $0.15 on qwen3.8-flash. The practical consequence is that answer length is the expensive half of a request, so an instruction to write short replies saves more than trimming a pasted document.
Yes, where the model supports it. Alibaba's sentence is that if a model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. That turns a $3.94 overnight job into $1.97. The one restriction worth knowing: the batch discount and the context cache discount cannot be applied at the same time.
It can, but Alibaba does not publish the rate for the qwen3.8 models. Caching discounts input tokens only, and it cannot be combined with the batch discount. The cache page gives 125% for writing a cache and 10% for reading one as its general example, then says qwen3.8-max, qwen3.8-max-0902, qwen3.8-flash and qwen3.8-2.4t-a95b do not follow that rule and points to the console instead. This page leaves the number out rather than quote the general example as if it applied.
For the first million tokens of each model it does, because those are free. The quota exists in the Singapore region with the service deployment scope set to International, is counted per model, covers real-time inference only, and expires after 90 days. Once it is gone, calls are billed at the rates above unless you have turned on Free Quota Only, which stops the service instead.
qwen3.8-flash, on published numbers, by a wide margin: the same monthly volume costs $0.92 on flash against $12.00 on max. Nothing in the price list says one model is better than the other; it prices both and leaves the choice to you. Model selection moves the bill about thirteen times, which is more than any of the other levers on this page.

How these numbers were checked

The two rates and the batch sentence come from Alibaba Cloud's model inference pricing page, which is linked below and carries its own last-updated stamp; when we read it on October 6, 2026, that stamp said October 6. The cache rule comes from the Context Cache page, and the free-quota conditions from the new-users page. Every figure on this page is one of those, or arithmetic done on them.

The arithmetic is multiplication, and you are meant to redo it: monthly tokens in millions, times the rate per million. Where a total would land between cents we have stayed with volumes that do not, so nothing here is rounded. If a number here ever disagrees with Alibaba's page, their page wins, and a line to support@llmomics.com gets it fixed.