Requirements and speeds checked October 5, 2026

Strata LLM

Every number below opens the page it was read from: Niko1221's repository for the hardware and the speeds, Alibaba Cloud's price list for the hosted rate.

Strata LLM is an open-source engine that runs a 125-billion-parameter model on a desktop PC instead of a server. It isn't ours, and it was built for exactly one model, Qwen3.8-Flash-Next, on the kind of machine most of us already own. This guide covers what it needs, how to install it, how to point your own apps at it, and how fast it goes. It is also the other end of the story this site usually tells, about what the APIs charge per token.

What Strata LLM is, and what it isn't

Strata is an inference engine: the program that loads a model into memory and turns your questions into answers. It comes from a developer who publishes as Niko1221 and sits on GitHub under the MIT licence. What makes it worth a look is the machine it aims at. A normal desktop with one gaming graphics card, rather than a server with hundreds of gigabytes of graphics memory.

The model underneath is Qwen3.8-Flash-Next, a mixture-of-experts model of roughly 125 billion parameters. That does not fit on a 12 GB card, so Strata spreads the work across the computer. The part every word needs stays on the graphics card, all 24,576 experts live in system memory, and the few the card does not hold are computed by the processor at the same time. Niko1221's own summary is more precise than any paraphrase: an engine built for exactly one model and exactly this kind of PC.

Two things it is not. It is not a runner you point at any model file, and it is not a hosted service, because nothing leaves your machine. What it does offer is an API on your own computer that speaks the same shapes as OpenAI's and Anthropic's, so the editors and scripts you already use can be aimed at your desktop instead of at a paid endpoint.

Hardware requirements for Strata LLM

The minimum is 12 GB of video memory or more, 32 GB of system memory at the very least, and about 80 GB of free disk. Everything below is taken from the project's own installation notes, and each figure links to the line it came from.

Graphics card
8 GB starts, but slowly
System memory
64 GB runs every size
Free disk
An NVMe SSD if you have one
Operating system
Ubuntu 22.04 and 24.04 get everything installed for you
Driver
NVIDIA. AMD uses the kernel's amdgpu driver on Linux
Processor
Any Intel or AMD desktop chip from about the last eight years

The supported NVIDIA cards run from the RTX 20 series through the RTX 50 series. AMD support is narrower: the RX 7900 XT and XTX, the RX 7800 XT and 7700 XT, the RX 9060 XT, the RX 9070 and 9070 XT, the Radeon AI PRO R9700, and the RX 6800 and 6900 series. Integrated Radeon graphics are listed as not supported.

The driver is the one thing you install yourself. The installer handles the rest on the first run, and when no ready-made engine matches your card it offers to install the build tools and compile one, which takes 20 to 40 minutes and happens once.

Which size of Strata LLM fits your PC

There is one model, compressed to four sizes and offered in four flavours. Smaller sizes answer faster, larger ones are a little smarter, and the installer reads your memory and recommends one. Press Enter and it takes its own advice.

Size Your RAM Speed Quality
Q2_0 48 GB Fastest Good
IQ2_XS 48 GB Fast Better. The project's recommended pick
IQ3_XXS 64 GB Slower Great
IQ3_S 64 GB Slowest Best. Matches the full model on the published tests
Coder (IQ1_M) 32 GB Fastest prompts Code only, roughly. 91% of the full model's SWE-bench score

The RAM column is the project's own rule of thumb: enough memory for the experts plus about 10 GB for Windows and everything else open. Verified October 5, 2026.

The four flavours

Coder is ISTA-DASLab's coding version. It keeps half of each layer's experts, chosen on code and agent data, which buys it a 32 GB footprint and the fastest prompt reading of any size, at the cost of being weaker outside code and in languages other than English. Swift 1.5 is a fine-tune by UkisAI that thinks for a shorter time before it answers, so the reply arrives sooner at about the same quality. Unsloth's UD-IQ4_XS is a roughly 4-bit build between IQ3_S and the largest quants in quality, a 94 GB download that wants 48 GB of memory or more. OrcaRouter's uncensored IQ3_XXS needs a manual packing step and is not in the installer's menu.

You can add a second size or flavour later by running SETUP.bat, and shared files are not downloaded twice.

How to install Strata LLM

Two roads, same destination. You can hand the whole job to an AI coding assistant, or do three steps yourself. Either way you end up with the app open in your browser at 127.0.0.1:8080.

Let an assistant do it

If you already use Claude Code, Cursor, Codex or GitHub Copilot, paste this into it. It checks your graphics card, memory and disk, picks the size that fits, installs, starts the model, and tells you how to connect your apps.

Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

Windows, by hand

  1. Download the project as a zip from the repository and unzip it somewhere with room for it, or clone it with git.
  2. Double-click START-HERE.bat.
  3. Answer three questions: which model and size, how much context, and whether it should read pictures. Press Enter at each one to take the recommended answer.

Then it downloads the model, around 70 GB, and starts it. If the download stops, run the same file again and it picks up where it left off.

Linux, by hand

Run ./setup.sh in the Strata folder. The questions and the install are the same, and it opens the same address.

What the first start looks like

Your PC can become slow or stop responding for one to three minutes while Strata loads 35 to 55 GB into memory and locks part of it for the graphics card. That is normal on the first run, and the window tells you what it is doing. Wait it out and don't close it; later starts take 30 to 90 seconds. Still frozen after ten minutes? Restart, close your browsers, and try a smaller size.

AMD cards

The commands are the same. Setup finds the Radeon card and chooses the AMD engine by itself, installs ROCm into a private folder on Linux without asking for a password, and compiles the engine for your card. Two things differ for now: pictures are read by the processor rather than the GPU, and the --calibrate tuning pass is NVIDIA-only.

Updating

UPDATE.bat and ./update.sh bring in a new version without starting the model, which is what you want when the graphics card is busy. Model files are left alone, and a running engine cannot be replaced underneath itself, so close the model's window first.

Running Strata LLM and pointing your apps at it

Open 127.0.0.1:8080 in a browser and you get the app that ships with the engine: a chat, a live monitor showing what the model and your hardware are doing, and an about tab with the addresses.

The more useful part is what sits behind it. Strata serves an API on your own machine that mirrors OpenAI's and Anthropic's, so anything that lets you set a base URL can use it. In an OpenAI-compatible client that is http://127.0.0.1:8080/v1, and any API key and any model name are accepted. Apps that speak Anthropic's API use /v1/messages instead, and Claude Code takes ANTHROPIC_BASE_URL=http://127.0.0.1:8080.

Thinking effort has four settings, off through high, and off is the fastest. The first message of a conversation is read in full and is the slow part, roughly a minute for every 30,000 tokens; follow-ups start in seconds. To reach it from another device, start it with --host 0.0.0.0 --api-key and a secret of your own, and treat that key as mandatory once the port is reachable from anywhere but your own machine.

Strata LLM speed calculator

Niko1221 has published measured speeds for two machines and an estimate for a third. There is no published model of how some other card would behave, so this calculator does not invent one. Tell it what you have and it will say whether the size you picked fits, and give you the measured figure from the closest machine the project actually tested.

What the project measured

Size RTX 5070, 12 GB
answers / prompt
RX 9070 XT, 16 GB
answers / prompt
Q2_0 94 / 2,650 60 / 1,160
IQ2_XS 79 / 2,090 52 / 1,110
IQ3_XXS 62 / 1,750 —
IQ3_S 53 / 1,620 —
Coder 55 / 2,180 44 / 1,420

The first number is how fast the reply appears in a short chat; the second is how fast a 32,000-token document or code file is read. A token is about three quarters of a word, so 60 tokens per second is faster than most people read. Both columns are the project's own measurements.

One more figure, and its limit. The README says a card with more memory is faster, and estimates an RTX 3090 with 24 GB at about 100 to 140 tokens per second. It does not say which size that figure is for, so it stays out of the table and out of the calculator's arithmetic.

Work it out for your PC

Will it run? Fits
Memory this size needs—
Verdict—
Answers, tokens per second Measured
Writes answers—
Reads your prompt—
Where the figure comes from—

A dash means the project has published no figure for that combination, and we would rather show a dash than a number nobody measured. The fit check uses the memory rule in the sizes table above. Graphics memory is not part of that rule, with one exception the installer handles by itself, when there is too little memory for the normal mode and it maps the experts from the disk instead.

Running it yourself, against paying per token

Alibaba's price list sells a hosted model called qwen3.8-flash. Strata runs a model called Qwen3.8-Flash-Next, from the project's own repository. Alibaba's pages never use that second name, so we do not claim they are the same model, and their prices do not go in one table. What follows is two ways to get similar work done, each with its own source.

On someone else's machine
Per million tokens, International list, requests up to a million tokens
On your own machine
No bill per token
A card from 12 GB, 32 GB of memory, about 80 GB of disk, and the power it draws

The hosted rate was read and checked on October 6, 2026. The hardware figures above are the ones checked at the top of this page.

The two bills are shaped differently. Hosted, you pay for what you send: your month's tokens divided by a million, times the two rates above. Local, the model is already yours, so the question stops being the price of a token and becomes whether the machine you own can run it, and what it draws from the wall. Put a month of your own traffic next to the price of a card and the answer usually turns on how much you send, not on which way is cheaper per word.

Two limits on this. The speed figures further up are the project's own measurements on three machines; nobody has published how these engines hold up under a day-long load, so we do not extrapolate one. And the electricity is yours to fill in, your rate and your hours. We print no figure for it, because that would be a guess at your bill rather than a number somebody measured.

Strata LLM: common questions

An open-source inference engine from the developer Niko1221. It runs one large model, Qwen3.8-Flash-Next, on a normal desktop instead of a server, and it serves an OpenAI- and Anthropic-compatible API on your own machine. The code sits on GitHub under the MIT licence.
One model family: Qwen3.8-Flash-Next, in four sizes (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S) and four flavours (the original, the Coder version, the Swift 1.5 fine-tune, and Unsloth's roughly 4-bit builds). It does not load other models.
A graphics card with 12 GB of memory or more, 32 GB of system memory at the very least, about 80 GB of free disk, and a driver no older than version 580 on NVIDIA. The installer sets up Python, the engine and the model by itself.
32 GB runs the Coder size. 48 GB runs Q2_0 or IQ2_XS. 64 GB runs every size, though IQ3_S wants little else open.
On Niko1221's RTX 5070 with 64 GB of memory, replies run from 53 to 94 tokens per second depending on the size, and a 32,000-token prompt is read at 1,620 to 2,650 tokens per second. A token is about three quarters of a word.
No. The only thing you install yourself is a graphics driver. On the first run the installer fetches Python if you have none, a private environment for it, NVIDIA's CUDA libraries, a ready-made engine for your card, and the model.
Point them at http://127.0.0.1:8080/v1 as an OpenAI-compatible provider. Apps that speak Anthropic's API use .../v1/messages, and Claude Code takes ANTHROPIC_BASE_URL=http://127.0.0.1:8080. Any API key and any model name are accepted.
The engine is free under the MIT licence; the model files come from other teams and carry their own licences. Your cost is the machine and the electricity, weighed against a per-token bill. If you already own a card with 12 GB and enough memory, local messages cost nothing per message; if you would have to buy the machine, paying per token usually comes out cheaper until you send a lot of them.
No. The model runs on your machine and the chat page is served from it, so your conversations never go to someone else's server.

How we check these Strata LLM numbers

Every figure on this page was read out of the source it names: Niko1221's own repository for the measured speeds, the hardware and the model sizes, and Alibaba Cloud's price list for the one hosted rate. Each number opens the page it came from. We have not run Strata ourselves, and where a figure is the project's own measurement we name the machine it was measured on.

Strata moves fast: it went from its first commit to version 0.1.39 inside eleven days, and speeds, download sizes and memory thresholds move with it. That is what the date at the top is for, and the honest cadence here is a re-read once a month. If a number disagrees with the repository, the repository is right. Nothing on this page is an advertisement, and if you find a figure that no longer matches its source, write to support@llmomics.com.