Run Qwen-Image-2.1-Turbo Locally on Windows 11, Mac and Kali Linux: 8-Step Install Guide (ComfyUI, GGUF, Python), the One Setting Everyone Misses, and the License Catch
Qwen-Image-2.1-Turbo is the fast version of Alibaba's open image model, released on October 9, 2026. It draws and edits pictures in 8 denoising steps instead of the 40 the base model's examples use, on the same 7-billion-parameter image generator, with the same 2048-pixel output, the same text rendering and the same editing with reference photos. You can run it locally on Windows 11, Mac and Kali Linux, free, through ComfyUI, a GGUF file as small as 3.2 GB, Python, or an Apple-silicon build. Within a day of release there were 23 community quantizations of it. Here is the surprise that catches almost everyone: setting "8 steps" in your app is not the same as running Turbo. The model carries its own eight sampling values inside the download, and ComfyUI, mflux and even the official Python library ignore them unless you hand them over. And here is the catch: it ships under the Qwen Research License, research and evaluation only, so the poster for your shop window needs a different model or the paid API at $0.016 an image.
Jake runs a phone repair shop. He wanted a poster for the window, a few pictures for the shop's social pages, and he had read that the new Qwen model makes a 2K image "in seconds" on an ordinary graphics card. His friend Ethan, a developer, came over on Friday evening with a laptop, a 12 GB graphics card and a plan. This page is what they learned: what Turbo actually is, which of the many files to download, how much memory each one needs, how to install it on Windows 11, Mac and Kali Linux, the one setting that makes the difference between "Turbo" and "an 8-step blur," what it costs in the cloud, and why Jake's poster ended up being made with a different model entirely.
New to running AI on your own machine? Our plain-English series on running AI locally explains models, parameters and file formats in about ten minutes. You can follow this page without it; every step starts from zero. If you already run the base model, our Qwen-Image-2.1 install guide covers the parts Turbo shares, and this page only repeats what changed.
What Qwen-Image-2.1-Turbo is, in plain English
An image model like this one does not draw a picture in one go. It starts with a square of random noise and cleans it up in rounds, and each round is a "denoising step." The base Qwen-Image-2.1, released September 20, 2026, uses 40 of those rounds in Alibaba's own examples. Every round runs the whole 7-billion-parameter generator once, so 40 rounds is 40 passes through a large model, and that is where your waiting time goes.
Turbo is an "accelerated checkpoint": the same architecture, retrained so that it reaches a finished picture in 8 rounds. Alibaba announced it on October 9, 2026, alongside the paid Pro and Turbo versions of the model in its cloud API. In the words of the Qwen team's own announcement, "fewer steps does not mean lower quality: it still generates strong 2K images." That is the maker's claim, and we will come back to how much of it has been measured.
Three technical details matter for what follows, and they are the whole reason this page exists rather than a one-line "set steps to 8":
- The schedule lives in the weights. The download's configuration file carries eight exact sampling values, called sigmas. They are not evenly spaced. The official Python library reads them automatically; most other software does not.
- Guidance is off. The model runs with CFG (classifier-free guidance, the "how hard to obey the prompt" dial) at 1, which means no guidance pass at all. That halves the work per step and means a negative prompt does nothing.
- It caches the prompt. With
use_kv_cacheswitched on, the model works out your text prompt and reference images once and reuses that work in every step, which matters most when you edit with several photos.
Everything else is inherited from the base model: native 2048 by 2048 output, legible text inside pictures, editing with up to ten reference images, and real transparent PNGs with an alpha channel. The 2.1 line replaced the older Qwen-Image models' 20B design with a lighter 7B one, which is what made a Turbo version practical on home hardware.
Jake: "So it's the same model, just lazier?"
Ethan: "Same model, taught to finish in eight brushstrokes instead of forty. The trick is that it only knows how to do that in a very particular order. Give it eight evenly spaced strokes and it has never practiced that."
Turbo vs Qwen-Image-2.1 vs the Turbo LoRA: which file is which
Searches for this model are full of confusion between the Qwen image models, and the Turbo release adds one more option. Here is the whole family in one table:
| Name | What it is | Steps | Weights | License |
|---|---|---|---|---|
| Qwen-Image-2.1 | The base model, September 20, 2026. Generation and editing in one checkpoint. | 40 in the official examples | Open | Qwen Research License |
| Qwen-Image-2.1-Turbo | The 8-step checkpoint, October 9, 2026. This page. | 8, fixed schedule | Open | Qwen Research License |
| Turbo LoRA | Comfy-Org's 0.91 GB extraction of the difference between Turbo and the base model. Apply it to the base weights you already have. | 8 | Open | Qwen Research License |
| Qwen-Image-2.1-Pro | A larger hosted model in Alibaba's API only. | n/a | Closed | $0.04 per image |
| Qwen-Image 2512 / Edit 2511 | The older 20B generation and its editing sibling. Different architecture, different files. | varies | Open | Apache 2.0 |
| Community "turbo" LoRAs | Third-party 4-to-8-step adapters that appeared before the official Turbo (Viggle, Pruna, Turbo8 and others). | 4 to 8 | Open | Varies; the base license still applies |
Two practical rules fall out of that table. If you have never installed any Qwen image model, download the full Turbo checkpoint and skip the base model entirely; it generates and edits on its own. If you already run the base model, the 0.91 GB LoRA is the cheapest upgrade, and you keep the ability to switch the LoRA off for a slow, careful 40-step render when a picture deserves it.
One more trap: the official repository is still 32.5 GB even though the generator is "7B." The generator itself is 14.2 GB in full precision, and sitting next to it is the text encoder, a Qwen3-VL 8B model that reads your prompt and reference photos. The brain is bigger than the painter. Every route below reuses the encoder and VAE files from the base model, so if you installed Qwen-Image-2.1 before, the only new download is the 7.26 GB or 14.23 GB Turbo generator.
The license catch: research and evaluation only
The weights are published under the Qwen Research License Agreement, dated September 20, 2026, the same license as the base model. Its first section defines "non-commercial" as "for research or evaluation purposes only," and the grant of rights is "for non-commercial purposes only." Any commercial use needs "a separate commercial license" from Alibaba, requested at the address in the license file. There is no revenue threshold and no small-business exception. Earlier Qwen image models shipped under Apache 2.0; the 2.1 line changed that, and Turbo inherits it.
The license also asks that anything you build on the weights carries a "Built with Qwen" notice, that you do not use "Qwen" as the primary name of a derivative, and that a breach lets Alibaba terminate the license, at which point you must delete the model.
Jake: "It's a poster for my own window. Does that count?"
Ethan: "A poster that sells repairs is commercial use. Evaluating the model on your laptop to decide whether it is any good is exactly what the license allows. So we test with Turbo tonight, and if you like what it does, the poster gets made with a model that lets you, or with the paid API."
Those alternatives exist and are good. Z-Image-Turbo, also from Alibaba's Tongyi lab, is a 6B model that finishes in 8 forward passes and is licensed Apache 2.0, which permits commercial use. FLUX.2 [klein] 4B from Black Forest Labs is a 4-step, 4-billion-parameter model, also Apache 2.0. Neither matches Qwen-Image-2.1's editing with ten reference images or its transparent output, but for a poster they are the honest choice. Alibaba's hosted API is the other route, and we price it below; note that paying per image through the API is a separate agreement from the research license on the downloaded weights.
Qwen-Image-2.1-Turbo hardware: VRAM, RAM and disk for every file
Turbo changes the time per image, not the memory per image. Eight steps on a 7B generator need the same memory as forty steps; they just finish sooner. So the memory rules from the base model carry over exactly, and the question is only which precision of each file to download. Here is every official and semi-official file, with the exact sizes from the repositories:
| File | Source | Size | Runs in | Pick it when |
|---|---|---|---|---|
qwen_image_2.1_turbo_bf16.safetensors | Comfy-Org | 14.23 GB | ComfyUI | 24 GB card; you want the reference output |
qwen_image_2.1_turbo_int8_convrot.safetensors | Comfy-Org | 7.26 GB | ComfyUI | The default first install; 12 to 16 GB cards |
qwen_image_2.1_turbo_lora_avg_rank_178_bf16.safetensors | Comfy-Org | 0.91 GB | ComfyUI, on the base model | You already have Qwen-Image-2.1 installed |
Qwen-Image-2.1-Turbo-INT8-ConvRot.safetensors | Unsloth | 7.26 GB | Python (Diffusers) | Coders on 16 GB or less; closest to full quality of the small files |
Qwen-Image-2.1-Turbo-FP8.safetensors | Unsloth | 7.12 GB | Python (Diffusers) | Only if your hardware runs FP8 faster than INT8 |
qwen_image_2.1_turbo_Q4_K_M.gguf (and five siblings) | Community (Abiray and others) | 3.19 to 7.59 GB | ComfyUI with the GGUF loader | 6 to 12 GB cards |
qwen3vl_8b_int8_convrot.safetensors (text encoder) | Comfy-Org | 9.35 GB | ComfyUI | The default encoder; 17.53 GB in full precision, 6.31 GB as W4A8 |
qwen_image_2.1_vae_bf16.safetensors | Comfy-Org | 0.68 GB | ComfyUI, llama-style tools | Always; it is not interchangeable with older Qwen or Wan VAEs |
Unsloth measured how far each small file drifts from the full-precision original, using LPIPS, a perceptual difference score where lower is closer. The INT8 file with "ConvRot" (a rotation trick applied before rounding) scored 0.037, plain INT8 0.089, and FP8 0.125. So the 7.26 GB INT8-ConvRot file is both the one Comfy-Org ships as default and the most faithful of the small ones; the slightly smaller FP8 file is the least faithful. Smaller is not better here.
For total memory, Unsloth's guidance for the base model applies unchanged: a 12 to 16 GB graphics card starts with a GGUF Q4_K_M at 1024 by 1024; a 24 GB card runs the INT8 or FP8 files comfortably; a 6 GB card can run the FP8 file with CPU offload at "under 2x slower"; a machine with no usable GPU and 12 to 16 GB of RAM can run a GGUF on the processor with a quantized text encoder, slowly. Keep image dimensions divisible by 32. If you are deciding what to buy, our laptop guide for local AI explains why the graphics card's memory, not the processor, decides everything here.
Disk: the INT8 ComfyUI set (Turbo generator, INT8 encoder, VAE) is about 17.3 GB. The full-precision set is about 32.4 GB. The GGUF Q4_K_M set with the W4A8 encoder is about 11.2 GB. Add room for ComfyUI itself and for 2048 by 2048 PNGs, which are not small.
How to run Qwen-Image-2.1-Turbo in ComfyUI on Windows 11 (step by step)
ComfyUI has supported Qwen-Image 2.1 natively since version 0.37.0 on September 21, 2026, and Turbo uses the identical architecture, so it loads with the same nodes and the same template. The official Comfy-Org repackage already includes the Turbo files. If ComfyUI is new to you, our ComfyUI download and install guide covers Desktop, Portable and manual installs, the Manager, and the folder layout; here we assume ComfyUI is installed and start from the first launch.
- Update ComfyUI to 0.37.0 or newer. The current release is 0.39.0 (October 5), which also fixes a crash when choosing the int8 or int4 cache with Qwen Image 2.1. In Comfy Desktop, use the update prompt; in Portable, run the update script in the
updatefolder. - Download three files from the Comfy-Org "Qwen-Image-2.1" repository on Hugging Face and put each in its folder inside
ComfyUI/models/:diffusion_models/qwen_image_2.1_turbo_int8_convrot.safetensors(7.26 GB) intomodels/diffusion_models/text_encoders/qwen3vl_8b_int8_convrot.safetensors(9.35 GB) intomodels/text_encoders/vae/qwen_image_2.1_vae_bf16.safetensors(0.68 GB) intomodels/vae/
- Open the template. In ComfyUI, open Workflow, then Browse Templates, and pick "Qwen Image 2.1: Text to Image." For editing, pick "Qwen Image 2.1: Image Edit." There is also a "Remove Background" template that uses the transparent output.
- Swap the model. In the Load Diffusion Model node, choose the Turbo file instead of the base one. Leave the encoder and VAE nodes as they are.
- Set the sampler. In the KSampler node: steps 8, CFG 1.0, sampler euler, scheduler simple. Clear the negative prompt; at CFG 1 it is ignored anyway, and leaving it filled only wastes encoder time. Keep the width and height at the template's 2048 values or drop to 1024 by 1024 for a first test on a smaller card.
- Queue it. The first run loads 17 GB of files and is slow. The second run is the real one.
That gives you a working Turbo in ComfyUI, and for many people it is enough. It is not quite the Turbo that Alibaba evaluated, and the next section explains why and how to close the gap in two extra nodes.
The one setting that matters: Turbo's own eight sigma values
This is the part most guides skip, and it is the single most important thing on this page. The Turbo download's model_index.json contains this line:
"sample_sigmas": [1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568]
Those eight numbers are the noise levels at which the model was trained to take its eight steps, and they are nothing like evenly spaced. Five of the eight steps happen while the picture is still more than 84 percent noise; the last step jumps from 41 percent noise straight to a finished image. The model learned to do big, confident work in that final jump. A "simple" or "normal" scheduler set to 8 steps spreads the steps out differently, so you are asking Turbo to take jumps it never rehearsed. It still produces a picture, often a decent one; it is just not the model Alibaba tested, and on hard prompts (small text, hands, fine patterns) the difference shows.
Alibaba's model card is blunt about it: the saved schedule is used by default in its own library, setting num_inference_steps on its own "does not change it," and other schedules "have not been evaluated for this checkpoint."
Here is how to use the real schedule in each tool:
- ComfyUI: replace the KSampler with the custom sampling nodes. Add a ManualSigmas node (it is marked experimental; search "define sigmas"), and type the eight values followed by a zero:
1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568, 0. Add a KSamplerSelect node set to euler, and a SamplerCustom node with CFG 1.0; wire the sigmas and the sampler into it along with the model, the conditioning and the empty latent. Nine values give eight steps, because each step runs from one noise level to the next. - Python with Diffusers: nothing to do, provided your Diffusers is new enough. Details in the Python section.
- mflux on Mac: the community build's notes say stock mflux does not yet read the saved grid, so the Apple-silicon route currently runs an approximation.
- GGUF in ComfyUI: the same ManualSigmas trick works with the GGUF loader; the schedule belongs to the sampler, not the file format.
Jake: "Eight numbers. That's the whole secret?"
Ethan: "That's the whole secret. The weights are the painter; those numbers are the order of the brushstrokes the painter was taught. Every app lets you set how many strokes. Almost none of them ask which ones."
The GGUF route: Turbo on a 6 to 12 GB graphics card
GGUF is the file format that made local chat models practical on small machines, and ComfyUI can load image models in it through the ComfyUI-GGUF custom node by city96. Qwen has not published GGUF files for Turbo; within hours of release, community members converted them, and the most complete set at the time of writing is the "Abiray" repository, with six sizes. Because these are third-party conversions, treat the memory ranges below as the uploader's estimates, not Alibaba's.
| File | Size | Uploader's VRAM guidance | Notes |
|---|---|---|---|
| Q8_0 | 7.59 GB | 12 to 16 GB | Closest to the original; no real reason to pick it over the official INT8 file |
| Q6_K | 5.88 GB | 10 to 12 GB | The sweet spot for a 12 GB card at 2048 by 2048 |
| Q5_K_M | 5.01 GB | 8 to 12 GB | Good texture detail |
| Q4_K_M | 4.19 GB | 6 to 8 GB | The uploader's recommended balance |
| Q4_K_S | 4.06 GB | 6 to 8 GB | Slightly smaller, slightly rougher |
| Q3_K_M | 3.19 GB | 6 GB, laptops | Last resort; expect visible loss on text and fine detail |
Setup, assuming ComfyUI is already installed:
- Install ComfyUI-GGUF through the Manager (search "GGUF"), or clone
city96/ComfyUI-GGUFintoComfyUI/custom_nodes/and install its requirements. Update it if you installed it months ago; older versions do not recognize the Qwen Image architecture. - Put the chosen
qwen_image_2.1_turbo_Q*.gguffile inmodels/diffusion_models/(the oldermodels/unet/folder also works). - Use the same text encoder and VAE as the ComfyUI route:
qwen3vl_8b_int8_convrot.safetensors(or the 6.31 GBqwen3vl_8b_w4a8.safetensorsif memory is tight) andqwen_image_2.1_vae_bf16.safetensors. - Open the Qwen Image 2.1 template and replace the Load Diffusion Model node with the Unet Loader (GGUF) node. Pick your file.
- Sampler: euler, scheduler simple, 8 steps, CFG 1.0, empty negative prompt, or the ManualSigmas setup from the schedule section for the exact Turbo behavior.
The uploader suggests 4 to 8 steps and calls 6 "ideal." Be careful with that: 6 steps is a further approximation of an approximation, and fewer steps saves you perhaps a second. Start at 8 with the real sigmas and only go lower once you know what the model can do.
If generation falls back to the processor and takes minutes instead of seconds, the graphics card is usually not being used at all; our guide to why local AI ignores your GPU covers the driver and CUDA checks, which are the same for ComfyUI.
How to run Qwen-Image-2.1-Turbo on Kali Linux
Kali is Debian underneath, so everything above works, with two Kali-specific habits. First, Kali refuses system-wide pip install with the "externally-managed-environment" error; always work inside a virtual environment, as our guide to that error explains. Second, Kali ships a recent NVIDIA driver through its own packages, and the CUDA build of PyTorch must match it; check nvidia-smi runs before you install anything.
For ComfyUI on Kali, the manual install is the clean route:
sudo apt update && sudo apt install -y git python3-venv
git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv .venv && source .venv/bin/activate
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
python main.py
Then open http://127.0.0.1:8188 in a browser, drop the three files into the same models/ subfolders as on Windows, and follow the ComfyUI steps above. The Manager and the GGUF node install the same way. On an AMD card, use the ROCm PyTorch index instead of the CUDA one; on a laptop with no discrete GPU, use the GGUF route and expect minutes per image rather than seconds.
For Python without ComfyUI, the Diffusers route below works identically on Kali; just create the venv first.
Qwen-Image-2.1-Turbo on a Mac: the honest status
There is no official Apple-silicon path. Alibaba's code targets CUDA, and ComfyUI does run on Apple silicon through Metal, but neither Alibaba nor Comfy-Org publishes Mac guidance or timings for the 2.1 models. What exists, within a day of release, is a set of community conversions for Apple's MLX framework:
- Osaurus / mflux builds: 4-bit (9.9 GB), 6-bit (14 GB) and 8-bit (17 GB) versions packaged for the mflux library and the Osaurus app (version 0.25.21 or newer). Their defaults match the model: 8 steps, guidance 1.0, dimensions in multiples of 32. Their notes say plainly that stock mflux does not yet read the saved sampling grid, so these run an approximate schedule.
- MLX-Serve builds: 4-bit and 8-bit versions for serving the model over an API from a Mac.
Memory: 4-bit weights of 9.9 GB plus the text encoder and working memory means a 16 GB Mac is the realistic floor at 1024 by 1024, and 32 GB or more for comfortable 2K work. Nobody has published timings yet. If you only want to see what the model does, the hosted demo on Hugging Face and the API are quicker than fighting with a Mac this week, and we expect the Mac story to improve as mflux catches up with the sigma schedule.
The Diffusers route: Qwen-Image-2.1-Turbo in Python (Windows 11, Kali, Linux)
This is Alibaba's own route and the only one that reads the saved schedule automatically. You need an NVIDIA card; the examples move the whole pipeline to CUDA, and the full-precision model wants a 24 GB card unless you enable CPU offload or use Unsloth's INT8 file.
The model card tells you to install Diffusers from its GitHub source, because the feature that reads sample_sigmas was merged on October 5, 2026. We checked the next release: Diffusers 0.41.0, published October 6, already contains it, so an ordinary pip install is enough. The Transformers version matters too; the card requires 5.17.0 or newer.
python3 -m venv ~/turbo && source ~/turbo/bin/activate # Windows: turbo\Scripts\activate
pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install "diffusers>=0.41.0" "transformers>=5.17.0" accelerate pillow
Text to image, following the official example:
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1-Turbo", dtype=torch.bfloat16
)
pipe.to("cuda") # or: pipe.enable_model_cpu_offload() on a small card
image = pipe(
prompt="A hand-painted shop sign reading 'Jake's Phone Repair', warm evening light, 35mm photo",
width=1680, height=2512, # a 2:3 portrait; see the resolution table
use_kv_cache=True,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("sign.png")
Notice what is missing: there is no num_inference_steps and no guidance value. The pipeline reads the eight sigmas from the checkpoint and runs exactly eight steps at CFG 1. If you pass num_inference_steps=20 you still get the saved eight-step schedule; the only way to change the schedule is to pass an explicit sigmas=[...] list, and the card says any other list is untested.
Editing with a reference image uses the same pipeline; pass the picture as an RGBA image and describe the change:
from PIL import Image
ref = Image.open("shopfront.jpg").convert("RGBA")
out = pipe(
prompt="Replace the faded sign with a clean navy sign reading 'Jake's Phone Repair'",
image=[ref], # up to ten reference images
width=2048, height=2048,
use_kv_cache=True,
generator=torch.Generator("cuda").manual_seed(7),
).images[0]
out.save("shopfront-new.png")
Keep the output size on the model's grid. These are the resolutions Alibaba lists for Turbo:
| Aspect | Width by height |
|---|---|
| 1:1 | 2048 by 2048 |
| 4:3 / 3:4 | 2400 by 1792 / 1792 by 2400 |
| 3:2 / 2:3 | 2528 by 1696 / 1696 by 2528 |
| 16:9 / 9:16 | 2752 by 1536 / 1536 by 2752 |
For a card with 16 GB or less, swap the generator for Unsloth's INT8-ConvRot file. Unsloth's repository pairs it with the FP8 text encoder and the VAE from its base-model upload, and its card shows the loading code; the schedule still comes from the Turbo configuration, so the eight steps are preserved.
Using the Turbo LoRA on the base model instead
If the 14.23 GB base generator is already on your disk, Comfy-Org's 0.91 GB LoRA turns it into Turbo without a second large download. In ComfyUI, add a LoraLoaderModelOnly node between the Load Diffusion Model node and the sampler, pick qwen_image_2.1_turbo_lora_avg_rank_178_bf16.safetensors at strength 1.0, and use the same 8-step, CFG 1 settings (or the manual sigmas). The file name tells you what it is: an average-rank-178 extraction of the difference between the two checkpoints, which means it is close to, not identical with, the full Turbo model. Comfy-Org does not say which to prefer. Our rule: the LoRA for convenience and for switching between fast and slow renders, the full checkpoint when you want exactly what Alibaba shipped.
How much faster is Qwen-Image-2.1-Turbo really?
Alibaba publishes no timings, and no independent benchmark existed at the time of writing. Here is what the arithmetic says, and what it does not.
Against the base model's 40 steps, Turbo takes 8: five times fewer generator passes. Because CFG is 1, each step is also a single pass rather than the two that guidance needs, so against a base-model setup running with guidance the sampling work drops by roughly ten times. That is the part that scales with steps, and on a 24 GB card at 2K it is the bulk of the wait.
What does not shrink: the text encoder reads your prompt once (a few seconds for an 8B model), the VAE decodes the final 2K image once (a second or two, and memory-hungry), and the model files load from disk on the first run. On a small card using CPU offload, those fixed costs plus the shuffling of weights between RAM and the graphics card can take longer than the eight steps themselves, which is why Turbo feels dramatic on a 24 GB card and merely helpful on a 6 GB one.
Jake: "So on my laptop it's not 'five times faster'?"
Ethan: "On your laptop the painter was never the slow part. The painter got five times faster; the person carrying the canvas in and out of the room did not."
If you want an honest number for your own machine, generate the same prompt and seed with the base model at 40 steps and with Turbo at 8, both at 1024 by 1024, and time the second run of each (the first run includes loading). That comparison, on the same hardware, is the only speed figure worth repeating.
Is Turbo as good as Qwen-Image-2.1? What has and has not been measured
Alibaba's claim is that fewer steps "does not mean lower quality." Its evidence, so far, is a showcase of sample images on the model card. There is no Turbo row in any benchmark table, no human-preference study, and the company's own note that other schedules "have not been evaluated" tells you how narrow the tested path is. Every distilled fast model in this field has traded something, usually diversity between seeds and the finest detail, for speed. Expect the same here until someone measures it.
What has been measured is the cost of the small files, from Unsloth's comparison against full precision:
| File | Size | LPIPS vs full precision (lower is closer) |
|---|---|---|
| INT8-ConvRot | 7.26 GB | 0.037 |
| INT8 | 7.26 GB | 0.089 |
| FP8 | 7.12 GB | 0.125 |
The practical reading: with the INT8-ConvRot file, a 7.26 GB download gets you within a few percent of a 14.2 GB one. The GGUF files have no such measurement yet; the uploader's descriptions ("indistinguishable," "reference grade") are opinions, not numbers.
The test worth doing at home, for anyone who wants to judge it rather than take anyone's word: generate the same prompt at the same seed three ways, Turbo at 8 steps on the saved sigmas, Turbo at 8 steps on the "simple" scheduler, and the base model at 40 steps, all at 1024 by 1024. Put lettering in the prompt, because text is where fast models fail first, and look at that before anything else. One laptop and one prompt is not a benchmark, but it is the quickest way to see the schedule rule with your own eyes.
The hosted API: $0.016 per image, and what that buys compared with local
The same day the weights appeared, Alibaba made Turbo and Pro available in its Model Studio API. These are the published prices on the international (Singapore) deployment, from the pricing page updated October 9, 2026, with the hosted rivals beside them:
| Model | Price per output image | Free quota | 1,000 images |
|---|---|---|---|
| qwen-image-2.1-turbo | $0.016 | 10 images | $16 |
| qwen-image-2.1-pro | $0.04 | 10 images | $40 |
| qwen-image-plus (older model) | $0.03 | 100 images | $30 |
| z-image-turbo | $0.015 ($0.03 with prompt rewriting) | 100 images | $15 |
| Google Nano Banana 2.1 (Gemini API) | from $0.0336 per 1K image | none on the API | from $33.60 |
Input images for the 2.1 models are not billed; only successful outputs are. The China (Beijing) and "Global" deployments list Turbo at $0.014133 and Pro at $0.035333. Rate limits are per account and not on the public price page.
Two things the price table does not say. First, the API is how Jake's poster actually gets made: a paid API call is a commercial service with its own terms, separate from the research license on the downloaded weights. Second, at $0.016 an image, local generation pays for itself only in volume or in privacy. Ten test images cost nothing on the API's free quota; a thousand product photos a month cost $16; a workflow where you iterate fifty times on each picture is where your own graphics card wins. And nothing you generate at home is uploaded anywhere, which matters more for a customer's cracked phone than for a poster. For the hosted side in more depth, our Nano Banana 2.1 guide covers Google's free tier and prices, and our Gemini free-tier page explains why image models are the exception to it.
If you would rather rent a graphics card than buy one, Turbo's 8-step speed makes it a good fit for a cloud GPU you switch on for an hour; our guide to Hugging Face models on AWS covers the instance types and what an hour costs.
Qwen-Image-2.1-Turbo not working? Troubleshooting by symptom
The picture is blurry, washed out or has broken text at 8 steps
Almost always the schedule. Use the eight sigma values from the schedule section rather than a generic 8-step scheduler, keep CFG at exactly 1.0, and make sure you loaded the Turbo file, not the base model with steps set to 8. A base model at 8 steps produces exactly this kind of image.
"Unsupported architecture" or the GGUF will not load
Update the ComfyUI-GGUF node and ComfyUI itself; older versions predate the Qwen Image 2.1 architecture. In a manual install, git pull inside both folders and restart.
Out of memory on the first step
Drop to 1024 by 1024 for the test, switch the text encoder to the 6.31 GB W4A8 file, pick a smaller GGUF, keep batch size at 1, and close other GPU programs, including the browser's hardware acceleration if memory is tight. On the Python route, add pipe.enable_model_cpu_offload() and remove pipe.to("cuda").
Out of memory at the very end, after all eight steps
That is the VAE decoding a 2K image, not the generator. In ComfyUI, use the VAE Decode (Tiled) node; in Python, enable VAE tiling. The eight steps worked; the last second ran out of room.
Diffusers runs 20 steps when you asked for 8, or 8 when you asked for 20
Both are the schedule logic working as designed. With Diffusers 0.41.0 or newer the Turbo pipeline always uses the saved eight sigmas and ignores num_inference_steps. With an older Diffusers, the saved sigmas are ignored and you get whatever step count you asked for, on an untested schedule. Upgrade, and stop passing a step count.
The negative prompt changes nothing
Correct, and expected. At CFG 1 there is no guidance pass, so the negative prompt is never read. Describe what you want in the positive prompt instead; the model is good at following detailed instructions.
Each image takes minutes on a laptop with a graphics card
The card is probably not being used. Check that PyTorch sees it (torch.cuda.is_available() in Python, or the GPU line in ComfyUI's startup log). Our guide to local AI ignoring the GPU walks through drivers, CUDA versions and laptop power settings.
"Crash when selecting int8 or int4 cache"
A ComfyUI bug fixed in 0.39.0 on October 5, 2026. Update.
Colors look slightly off compared with a friend's output
Different quantizations give slightly different pictures: Unsloth's measurements put FP8 three times further from the original than INT8-ConvRot. Compare files, not just seeds. And make sure both of you use the Qwen-Image-2.1 VAE; the older Qwen Image and Wan VAEs load without error and decode wrongly.
Qwen-Image-2.1-Turbo: frequently asked questions
What is Qwen-Image-2.1-Turbo?
An accelerated version of Alibaba's open Qwen-Image-2.1 model, released October 9, 2026, that generates and edits images in 8 denoising steps instead of 40. It uses the same 7B generator, the same 2048-pixel output and the same editing with reference images.
Is Qwen-Image-2.1-Turbo free?
The weights are free to download and run on your own computer for research and evaluation under the Qwen Research License. Commercial use needs a separate license from Alibaba. The hosted API costs $0.016 per image with a 10-image free quota.
Can I use Qwen-Image-2.1-Turbo commercially?
Not under the downloaded weights' license, which allows research and evaluation only. For commercial images, use Alibaba's paid API, or an Apache 2.0 model such as Z-Image-Turbo (6B, 8 steps) or FLUX.2 klein 4B.
How much VRAM does Qwen-Image-2.1-Turbo need?
The same as the base model, because fewer steps do not reduce memory. A 24 GB card runs the full-precision file; 12 to 16 GB cards run the 7.26 GB INT8 file or a GGUF Q6_K; 6 to 8 GB cards use a GGUF Q4_K_M at 1024 by 1024, or FP8 with CPU offload.
How do I run Qwen-Image-2.1-Turbo in ComfyUI?
Update to ComfyUI 0.37.0 or newer, put the Comfy-Org Turbo file in models/diffusion_models, the Qwen3-VL 8B encoder in models/text_encoders and the 2.1 VAE in models/vae, open the Qwen Image 2.1 template, select the Turbo file, and set 8 steps with CFG 1.0. For exact results, use a ManualSigmas node with the model's eight sigma values.
What are the correct sigmas for Qwen-Image-2.1-Turbo?
1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568, taken from the model's configuration file. In ComfyUI's ManualSigmas node, add a final 0 to make eight steps. Diffusers 0.41.0 and newer read them automatically.
Why does setting 8 steps not give the same result as Turbo?
Because Turbo was trained on eight specific noise levels that are not evenly spaced; five of the eight steps happen above 84 percent noise. A generic 8-step scheduler uses different levels, which Alibaba says it has not evaluated. Use the saved sigmas.
Is there a Qwen-Image-2.1-Turbo GGUF?
Not from Qwen, but community conversions appeared within hours, in six sizes from a 3.19 GB Q3_K_M to a 7.59 GB Q8_0. They run in ComfyUI with the ComfyUI-GGUF custom node, alongside the standard text encoder and VAE files.
Does Qwen-Image-2.1-Turbo run on a Mac?
Only through community Apple-silicon builds (4-bit at 9.9 GB, 6-bit and 8-bit) for mflux and the Osaurus app. There is no official Mac path, and those builds do not yet use the model's saved schedule. Plan on a 16 GB Mac as the floor and 32 GB for 2K work.
How do I run Qwen-Image-2.1-Turbo on Kali Linux?
Inside a Python virtual environment, install ComfyUI manually with a CUDA build of PyTorch, place the three model files in the models folders, and use the ComfyUI steps on this page. The Diffusers route also works in a venv with diffusers 0.41.0 or newer.
Do I need the base model to use Turbo?
No. The Turbo checkpoint generates and edits on its own. If you already have the base model, the 0.91 GB Turbo LoRA from Comfy-Org turns it into Turbo without a second large download.
What is the difference between Qwen-Image-2.1-Turbo and the Turbo LoRA?
The full checkpoint is what Alibaba trained and tested. The LoRA is Comfy-Org's rank-178 extraction of the difference between Turbo and the base model, applied on top of the base weights; it is close to, not identical with, the full model, and lets you switch fast mode off.
How fast is Qwen-Image-2.1-Turbo?
Five times fewer generator passes than the base model's 40 steps, and no guidance pass. Alibaba has published no timings. On a 24 GB card the saving is dramatic; on a small card using CPU offload, loading and the text encoder dominate and the gain is smaller.
Is Qwen-Image-2.1-Turbo as good as Qwen-Image-2.1?
Alibaba says fewer steps do not lower quality, but has published only sample images, no benchmark. Distilled fast models usually give up some detail and seed-to-seed variety. Test the same prompt and seed on both before deciding.
Does the negative prompt work with Qwen-Image-2.1-Turbo?
No. Turbo runs with CFG 1, which means no guidance pass, so the negative prompt is never read. Put everything into the positive prompt.
What Diffusers version does Qwen-Image-2.1-Turbo need?
The card says to install from source; the feature it needs was merged October 5, 2026, and Diffusers 0.41.0, released October 6, includes it. Pair it with transformers 5.17.0 or newer.
Can Qwen-Image-2.1-Turbo edit images?
Yes. Pass up to ten reference images and describe the change; the official example edits a 2048 by 2048 image in the same eight steps. Prefix KV caching reuses the reference images across steps.
What does the Qwen-Image-2.1-Turbo API cost?
$0.016 per output image on Alibaba Cloud Model Studio's international deployment, with 10 free images; the Pro model is $0.04. Input images are not billed. The Beijing and Global deployments list Turbo at $0.014133.
By ten o'clock Jake had a folder of shop signs and a clear verdict: the model is fast, and the eight numbers matter more than anything else on the page. The poster for the window got made on Sunday with Z-Image-Turbo, under a license that lets a shop use it, and the Qwen test images stayed what the license says they are: an evaluation, and an honest one. The laptop never uploaded a thing.
If you keep one line from this page
Turbo is not "8 steps"; it is eight specific steps, and your app has to be told which ones.
Use the saved sigmas, keep CFG at 1, pick the INT8-ConvRot file, and read the license before the poster goes in the window.
Revision note. Written October 10, 2026, the day after Alibaba released Qwen-Image-2.1-Turbo. "Measure twice, cut once," says the old carpenter's proverb; here it is eight numbers once, and every picture after that comes out right.
