Free-quota terms checked October 6, 2026

Is Qwen Free?

Short answer first, then the three conditions that decide whether it is true for you.

Qwen is Alibaba's model family, and you can use it without paying for a while. The free part is the new-account quota on Model Studio, and it arrives with conditions that most pages leave out. This page walks through them, says what the quota does not cover, explains what happens on the day it runs out, and ends with the one route that never produces a bill at all. Paid rates are on our Qwen pricing page, and running a Qwen model on hardware you own is what the Strata guide is for.

Yes, for 90 days, in one region, per model

Opening a Model Studio account gives you a pool of free tokens rather than a bill, and the figure people repeat is a million tokens per model. That number is real, but it is only half of the sentence. The quota exists in the Singapore region and nowhere else, it is counted separately for every model, and it expires whether or not you use it. Miss any one of those and the free part stops being true.

What the free quota buys is not a different class of service. It is the same model, on the same machines, answering at the same speed as a paid call. Where the two differ is at the edges: the quota runs out, it does not cover every kind of job, and no amount of topping up your account raises the ceiling on how fast you can send.

  Free quota Paid calls
Where it works Singapore region only. The region you are billed in.
Answer speed Same infrastructure, same speed. The same as the free quota.
Requests per minute Your account's limit. Identical ceiling. A top-up does not raise it.
What it covers Real-time inference only. Real-time inference and batch jobs.
When it ends At 90 days, or when the pool empties. When your balance runs out.

Every cell above comes from Alibaba's free-quota and rate-limit pages; both are linked at the foot of this page.

The three things the free quota depends on

Alibaba states all three in a single note, which is worth reading exactly as written, because each one is a way for the offer to fail quietly.

One region
Singapore
The region has to be Singapore and the service deployment scope has to be International. Anything else, and there is no free tier to spend.
One pool per model
Typically 1,000,000 tokens
The documentation hedges with "typically", so treat it as the usual figure. Quota cannot move between models, and a dated snapshot counts as a model of its own.
One deadline
90 days
Dated from whichever came last: activation, the model's release, or approval of your request for it. Standing still does not pause that clock.

Three smaller rules are easy to miss and expensive to discover late. Input and output tokens come out of the same pool, so a million is not a million each. Whatever is left on day 90 is voided, and a second account does not reset the gift. If your team uses RAM users under one Alibaba Cloud account, they all spend from the same pool.

What the free quota does not pay for

Real-time calls are covered, and that is the end of the list. Alibaba's warning section names the jobs that bill normally from the first second: batch invocations, fine-tuning, model deployment, custom models, PAI-DSW instances, and OSS storage or request fees. If your plan was to spend the free million on a fine-tune, that plan does not work.

One detail saves money for people with several keys. Deductions run in a fixed order, free quota first, then any resource plan, then a savings plan, and pay-as-you-go last. A general API key draws on the free quota; the dedicated keys sold with a Token Plan or Coding Plan do not touch it, so the quota sits there until a general key spends it.

Are free calls slower?

No. Alibaba's rate-limit page answers this in as many words: free quota and paid calls use the same infrastructure, and the same model generates at the same speed either way. Model response speed is unrelated to whether you pay. The quota controls one thing, whether a request is accepted.

The ceiling on how fast you can send is set at the account level, and it is independent of billing, so paying does not lift it. A top-up only keeps the account from being suspended over an unpaid balance; the documentation sends anyone who needs a higher limit to their business manager instead. On the international list the qwen3.8 models are marked Dynamic Rate Limiting with no fixed numbers, while the Global scope rows list 30,000 requests and 5,000,000 tokens per minute.

Two error codes tell you which wall you hit. A 429 means the requests-per-minute or tokens-per-minute limit, and it usually clears inside a minute. A 403 carrying the code AllocationQuota.FreeTierOnly means the pool is empty and the service stopped, which is the switch described next.

What happens when the free quota runs out

It depends on one checkbox on your account. An account that has completed identity verification moves to pay-as-you-go billing automatically, so the next call is simply charged at the published rates. An account that has not been verified cannot keep calling at all until it verifies and adds funds, which is a hard stop rather than a bill.

There is a way to get the hard stop on purpose. Free Quota Only, which some users call worry-free mode, makes the service return an HTTP 403 with the code AllocationQuota.FreeTierOnly once the pool is empty instead of spending your balance. Alibaba suggests turning it on for anything you would rather have fail than bill, and warns against leaving it on in production, for the obvious reason.

Once you are paying, the arithmetic is the only thing left to look at. Two rates do most of the work, $2 in and $6 out per million tokens on qwen3.8-max against $0.15 and $0.47 on qwen3.8-flash, and the worked cost examples run a few monthly volumes through them.

Running Qwen without an account

There is a version of this that never expires and never bills. Alibaba publishes open-source Qwen weights, and its own price list carries a section for hosting them, including qwen3.8-27b at $0.50 in and $3 out per million tokens. Download those weights and run them on your own machine and the cost stops being a rate at all: it becomes the hardware, the electricity, and your afternoon.

That route has its own wall, and it is memory rather than money. Our Strata guide works through it for one 125-billion-parameter Qwen model, from the 32 GB of system memory at the bottom end to the sizes that want 96 GB, along with the speeds the project measured on the cards it tested.

Is Qwen free: common questions

No, part of it is, for a while. A new Model Studio account gets a free quota of typically 1,000,000 tokens per model, and it exists only in the Singapore region with the service deployment scope set to International. It is valid for 90 days, and when the pool is empty the calls are billed at the published rates unless you have turned on Free Quota Only, which stops the service instead.
Alibaba's documentation says typically 1,000,000 tokens per model, and it keeps the word typically, so treat it as the usual figure rather than a promise. Input and output tokens come out of that one pool, the quota cannot be moved from one model to another, and a dated snapshot of a model counts as a model of its own.
The two common reasons are written into the terms. The quota applies in the Singapore region only, with the service deployment scope set to International, so a model you call from another region has none. It also expires 90 days after the later of your activation date, the model's release date, or the approval date of your request for it, and whatever is left is voided at that point.
Alibaba's rate-limit page answers the speed half directly: free quota and paid calls use the same infrastructure, and the same model generates at the same speed either way. The ceiling on requests and tokens per minute is set at the account level and does not depend on billing, so adding funds does not raise it. A top-up only prevents suspension over an unpaid balance.
A verified account switches to pay-as-you-go automatically and the next call is charged at the published rates. An unverified account cannot keep calling the models at all until the owner completes verification and adds funds. If you want the service to stop before it spends your balance, turn on Free Quota Only, which returns an HTTP 403 with the code AllocationQuota.FreeTierOnly once the quota is gone.
No. The free quota covers charges for real-time inference only. Alibaba's warning lists what it does not cover: batch invocations, model fine-tuning, model deployment, custom models, PAI-DSW instances, and OSS storage or request fees. Those bill normally from the first call.
Yes. Alibaba publishes open-source Qwen weights, and its own price list carries a section for models that are hosted rather than called, including qwen3.8-27b at $0.50 in and $3 out per million tokens. Running those weights on your own machine turns the cost into hardware and electricity, which is what our Strata guide works through for one 125-billion-parameter Qwen model.

How we check what is free

Every claim on this page was read out of an Alibaba Cloud page and links back to it. The free-quota terms come from the new-users page, the speed and ceiling answers from the rate-limit page, and the rates quoted in passing from the model inference pricing page. Where the documentation uses a hedge such as "typically", we keep the hedge rather than tidy it away.

Free tiers change more often than prices do, and Alibaba's own page carries a transition note: accounts activated before September 15 that never filled in their account information can keep spending what is left of their quota, and after that they have to complete those details to continue on pay-as-you-go. If something here disagrees with Alibaba's page, their page wins. Write to support@llmomics.com and we will correct it.