The basics
What Strata LLM is, and what it isn't
Strata is an inference engine: the program that loads a model into memory and turns your questions into answers. It comes from a developer who publishes as Niko1221 and sits on GitHub under the MIT licence. What makes it worth a look is the machine it aims at. A normal desktop with one gaming graphics card, rather than a server with hundreds of gigabytes of graphics memory.
The model underneath is Qwen3.8-Flash-Next, a mixture-of-experts model of roughly 125 billion parameters. That does not fit on a 12 GB card, so Strata spreads the work across the computer. The part every word needs stays on the graphics card, all 24,576 experts live in system memory, and the few the card does not hold are computed by the processor at the same time. Niko1221's own summary is more precise than any paraphrase: an engine built for exactly one model and exactly this kind of PC.
Two things it is not. It is not a runner you point at any model file, and it is not a hosted service, because nothing leaves your machine. What it does offer is an API on your own computer that speaks the same shapes as OpenAI's and Anthropic's, so the editors and scripts you already use can be aimed at your desktop instead of at a paid endpoint.
Requirements
Hardware requirements for Strata LLM
The minimum is 12 GB of video memory or more, 32 GB of system memory at the very least, and about 80 GB of free disk. Everything below is taken from the project's own installation notes, and each figure links to the line it came from.
The supported NVIDIA cards run from the RTX 20 series through the RTX 50 series. AMD support is narrower: the RX 7900 XT and XTX, the RX 7800 XT and 7700 XT, the RX 9060 XT, the RX 9070 and 9070 XT, the Radeon AI PRO R9700, and the RX 6800 and 6900 series. Integrated Radeon graphics are listed as not supported.
The driver is the one thing you install yourself. The installer handles the rest on the first run, and when no ready-made engine matches your card it offers to install the build tools and compile one, which takes 20 to 40 minutes and happens once.
Sizes and versions
Which size of Strata LLM fits your PC
There is one model, compressed to four sizes and offered in four flavours. Smaller sizes answer faster, larger ones are a little smarter, and the installer reads your memory and recommends one. Press Enter and it takes its own advice.
| Size | Your RAM | Speed | Quality |
|---|---|---|---|
| Q2_0 | 48 GB | Fastest | Good |
| IQ2_XS | 48 GB | Fast | Better. The project's recommended pick |
| IQ3_XXS | 64 GB | Slower | Great |
| IQ3_S | 64 GB | Slowest | Best. Matches the full model on the published tests |
| Coder (IQ1_M) | 32 GB | Fastest prompts | Code only, roughly. 91% of the full model's SWE-bench score |
The RAM column is the project's own rule of thumb: enough memory for the experts plus about 10 GB for Windows and everything else open. Verified October 5, 2026.
The four flavours
Coder is ISTA-DASLab's coding version. It keeps half of each layer's experts, chosen on code and agent data, which buys it a 32 GB footprint and the fastest prompt reading of any size, at the cost of being weaker outside code and in languages other than English. Swift 1.5 is a fine-tune by UkisAI that thinks for a shorter time before it answers, so the reply arrives sooner at about the same quality. Unsloth's UD-IQ4_XS is a roughly 4-bit build between IQ3_S and the largest quants in quality, a 94 GB download that wants 48 GB of memory or more. OrcaRouter's uncensored IQ3_XXS needs a manual packing step and is not in the installer's menu.
You can add a second size or flavour later by running SETUP.bat, and shared files are not downloaded twice.
Install
How to install Strata LLM
Two roads, same destination. You can hand the whole job to an AI coding assistant, or do three steps yourself. Either way you end up with the app open in your browser at 127.0.0.1:8080.
Let an assistant do it
If you already use Claude Code, Cursor, Codex or GitHub Copilot, paste this into it. It checks your graphics card, memory and disk, picks the size that fits, installs, starts the model, and tells you how to connect your apps.
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
Windows, by hand
- Download the project as a zip from the repository and unzip it somewhere with room for it, or clone it with git.
- Double-click START-HERE.bat.
- Answer three questions: which model and size, how much context, and whether it should read pictures. Press Enter at each one to take the recommended answer.
Then it downloads the model, around 70 GB, and starts it. If the download stops, run the same file again and it picks up where it left off.
Linux, by hand
Run ./setup.sh in the Strata folder. The questions and the install are the same, and it opens the same address.
What the first start looks like
Your PC can become slow or stop responding for one to three minutes while Strata loads 35 to 55 GB into memory and locks part of it for the graphics card. That is normal on the first run, and the window tells you what it is doing. Wait it out and don't close it; later starts take 30 to 90 seconds. Still frozen after ten minutes? Restart, close your browsers, and try a smaller size.
AMD cards
The commands are the same. Setup finds the Radeon card and chooses the AMD engine by itself, installs ROCm into a private folder on Linux without asking for a password, and compiles the engine for your card. Two things differ for now: pictures are read by the processor rather than the GPU, and the --calibrate tuning pass is NVIDIA-only.
Updating
UPDATE.bat and ./update.sh bring in a new version without starting the model, which is what you want when the graphics card is busy. Model files are left alone, and a running engine cannot be replaced underneath itself, so close the model's window first.
Using it
Running Strata LLM and pointing your apps at it
Open 127.0.0.1:8080 in a browser and you get the app that ships with the engine: a chat, a live monitor showing what the model and your hardware are doing, and an about tab with the addresses.
The more useful part is what sits behind it. Strata serves an API on your own machine that mirrors OpenAI's and Anthropic's, so anything that lets you set a base URL can use it. In an OpenAI-compatible client that is http://127.0.0.1:8080/v1, and any API key and any model name are accepted. Apps that speak Anthropic's API use /v1/messages instead, and Claude Code takes ANTHROPIC_BASE_URL=http://127.0.0.1:8080.
Thinking effort has four settings, off through high, and off is the fastest. The first message of a conversation is read in full and is the slow part, roughly a minute for every 30,000 tokens; follow-ups start in seconds. To reach it from another device, start it with --host 0.0.0.0 --api-key and a secret of your own, and treat that key as mandatory once the port is reachable from anywhere but your own machine.
Speed calculator
Strata LLM speed calculator
Niko1221 has published measured speeds for two machines and an estimate for a third. There is no published model of how some other card would behave, so this calculator does not invent one. Tell it what you have and it will say whether the size you picked fits, and give you the measured figure from the closest machine the project actually tested.
What the project measured
| Size | RTX 5070, 12 GB answers / prompt |
RX 9070 XT, 16 GB answers / prompt |
|---|---|---|
| Q2_0 | 94 / 2,650 | 60 / 1,160 |
| IQ2_XS | 79 / 2,090 | 52 / 1,110 |
| IQ3_XXS | 62 / 1,750 | — |
| IQ3_S | 53 / 1,620 | — |
| Coder | 55 / 2,180 | 44 / 1,420 |
The first number is how fast the reply appears in a short chat; the second is how fast a 32,000-token document or code file is read. A token is about three quarters of a word, so 60 tokens per second is faster than most people read. Both columns are the project's own measurements.
Work it out for your PC
A dash means the project has published no figure for that combination, and we would rather show a dash than a number nobody measured. The fit check uses the memory rule in the sizes table above. Graphics memory is not part of that rule, with one exception the installer handles by itself, when there is too little memory for the normal mode and it maps the experts from the disk instead.
Cost
Running it yourself, against paying per token
Alibaba's price list sells a hosted model called qwen3.8-flash. Strata runs a model called Qwen3.8-Flash-Next, from the project's own repository. Alibaba's pages never use that second name, so we do not claim they are the same model, and their prices do not go in one table. What follows is two ways to get similar work done, each with its own source.
The hosted rate was read and checked on October 6, 2026. The hardware figures above are the ones checked at the top of this page.
The two bills are shaped differently. Hosted, you pay for what you send: your month's tokens divided by a million, times the two rates above. Local, the model is already yours, so the question stops being the price of a token and becomes whether the machine you own can run it, and what it draws from the wall. Put a month of your own traffic next to the price of a card and the answer usually turns on how much you send, not on which way is cheaper per word.
Two limits on this. The speed figures further up are the project's own measurements on three machines; nobody has published how these engines hold up under a day-long load, so we do not extrapolate one. And the electricity is yours to fill in, your rate and your hours. We print no figure for it, because that would be a guess at your bill rather than a number somebody measured.
FAQ
Strata LLM: common questions
Method
How we check these Strata LLM numbers
Every figure on this page was read out of the source it names: Niko1221's own repository for the measured speeds, the hardware and the model sizes, and Alibaba Cloud's price list for the one hosted rate. Each number opens the page it came from. We have not run Strata ourselves, and where a figure is the project's own measurement we name the machine it was measured on.
Strata moves fast: it went from its first commit to version 0.1.39 inside eleven days, and speeds, download sizes and memory thresholds move with it. That is what the date at the top is for, and the honest cadence here is a re-read once a month. If a number disagrees with the repository, the repository is right. Nothing on this page is an advertisement, and if you find a figure that no longer matches its source, write to support@llmomics.com.