WAN Video GeneratorWAN Video Generator

Z-Image Turbo FP8: Run the 6B Model Locally on an 8GB GPU (2026 Guide)

Jacky Wangon 7 hours ago

My main machine has an 8GB card. For two years that has meant one thing: local image generation is somebody else's hobby. Every "runs locally" guide ended the same way — download the 12B checkpoint, load it, watch the VRAM counter hit the ceiling, read the OOM error, close ComfyUI, go back to a hosted tool.

So when a 6B model with an Apache 2.0 licence and an 8-step sampler showed up, I did not believe the VRAM claims. I ran the numbers myself: BF16, then FP8, then a GGUF quant on the same card, same prompt, same seed. This guide is what actually happened — the real VRAM ceilings, the generation times, what the quantised build costs you in image quality, the install that works on an 8GB card, and the honest answer to when you should stop fighting your hardware and use a hosted Z-Image generator instead.

TL;DR

  • Z-Image Turbo is a 6B single-stream diffusion transformer from Alibaba's Tongyi Lab, Apache 2.0, 8 steps, no CFG. It is the first model at this quality level that genuinely runs on a mainstream 8GB card.
  • The three practical builds and their real ceilings: BF16 needs about 14–16GB VRAM, FP8 fits in roughly 8GB, and community GGUF quants squeeze it onto 6GB with --lowvram. System RAM matters too — 16GB minimum for FP8, 32GB for the GGUF low-VRAM route.
  • On a 4090, a 1024px image lands in about 2–3 seconds. On an 8GB card in FP8, expect a few seconds more per image, which is still fast enough to iterate by eye rather than by clicking generate and walking away.
  • FP8 is not a quality downgrade you will notice. Side by side with BF16 at 1024px, the difference showed up in fine high-frequency texture, not in composition, prompt adherence or text.
  • My recommendation: run FP8 locally if you have 8GB and 16GB of system RAM, and keep a hosted generator for the days you need volume, a bigger resolution than your card allows, or a machine without a discrete GPU. You will want both.

Real Test: BF16 vs FP8 vs GGUF on One 8GB Card

Same prompt, same seed, same 1024×1024 output, three builds on an 8GB card with 32GB of system RAM.

Build Peak VRAM Loaded and ready 1024px image What changed in the output
BF16 (full precision) ~14–16GB needed — did not fit natively; ran only with heavy offloading and BLAS fallback 40–60s with offload 35–50s Reference quality; unusable speed on this card
FP8 (e4m3fn) ~8GB ~10s 5–8s Near-identical at a glance; slight softening in fine texture
GGUF Q8 quant ~7GB ~8s 6–9s Same as FP8 to my eye
GGUF Q4 quant ~6GB ~7s 5–7s Visible texture flattening in skin and foliage; prompt adherence still good

Four findings worth writing down.

1. The memory ceiling is real, and it is the whole story. BF16 on a 6B model means roughly 12GB of weights before you add the text encoder, the VAE and the activation buffers. That is why "it runs on 16GB" is the honest BF16 statement and why FP8 is the build that matters for anyone below a 12GB card. FP8 halves the weight footprint, and 8GB is genuinely enough.

2. System RAM is the constraint nobody mentions. FP8 needs 16GB or more of system RAM; the GGUF low-VRAM route wants 32GB because weights stream from system memory during sampling. If your card fits the model but your RAM does not, you will see the process stall rather than crash, which is a more confusing failure than an OOM.

3. Quality loss from quantisation is smaller than the community warned. In my runs, FP8 and Q8 were indistinguishable in composition, prompt following and text rendering. The difference appeared only when I looked at high-frequency texture — fabric weave, hair strands, foliage edges — and even then it took a side-by-side crop to see it. Q4 is where it becomes visible without effort. My advice: use FP8 as the default, use Q8 GGUF if you need the last gigabyte, and skip Q4 unless 6GB is all you have.

4. Speed is the reason this changes your workflow. At 5–8 seconds per image, prompt iteration becomes conversational. Compare that with the 30–60 second loop of the larger open models on the same hardware, where you start queuing batches instead of testing ideas. Eight steps with no CFG is what makes that possible.

One caveat for AMD users: there are credible reports of Z-Image Turbo FP8 loading and generating very slowly under ROCm (a 16GB RX 9070 XT on ROCm 7.2 was reported as unusually slow), so if you are on AMD, test before you plan around local generation.

If you want output today, start here: Launch Z-Image Now →

What Z-Image Turbo Actually Is

Z-Image Turbo
Developer Alibaba Tongyi Lab (Tongyi-MAI)
Architecture S3-DiT — single-stream diffusion transformer
Parameters 6B
Released Z-Image-Turbo, 26 November 2025; base Z-Image, 27 January 2026
Sampling 8 steps, no CFG
Licence Apache 2.0
Weights Hugging Face (Tongyi-MAI/Z-Image-Turbo) and ModelScope
Text-to-image Yes — the released Turbo checkpoint is a generation model

Two things follow from that spec sheet. First, Apache 2.0 is the permissive licence — commercial use is not the legal maze it is with some of the larger open models, though you should still read the licence file yourself rather than trusting a blog post, including this one. Second, the released Turbo and base checkpoints are text-to-image models. If you need character consistency across a set, that is a workflow problem rather than an input slot problem, and it is why I wrote a separate guide on keeping a character consistent with Z-Image Turbo rather than trying to solve it in the install.

ComfyUI Install: The 8GB Route

This is the sequence that worked for me. It assumes ComfyUI is already installed and working.

  1. Download the FP8 checkpoint. From the Tongyi-MAI/Z-Image-Turbo Hugging Face repo, the FP8 file is the one named as a fp8 safetensors checkpoint. The direct-download path is pip install -U huggingface_hub followed by HF_XET_HIGH_PERFORMANCE=1 hf download Tongyi-MAI/Z-Image-Turbo. The same checkpoints are mirrored on ModelScope, which is faster if you are in Asia.
  2. Place the files in ComfyUI's model folders. The checkpoint goes under models/diffusion_models, the text encoder under models/text_encoders, and the VAE under models/vae. This is the layout the Z-Image template workflows expect — if you drop the checkpoint into models/checkpoints and pick a generic loader, you will get a node error rather than an image.
  3. Load the Z-Image template workflow. Recent ComfyUI builds ship a template for this architecture. If yours does not, rebuild it from the model card's reference workflow: loader → text encoder → sampler with 8 steps and CFG 1 → VAE decode. Do not raise the step count expecting more detail; this model was trained for the low-step regime.
  4. Set your resolution as multiples of 16. 1024×1024 is the comfortable default, 1536-class sizes work if you have the headroom, and off-grid dimensions produce visible repetition artifacts. If you are on 8GB, stay at 1024 and upscale afterwards rather than generating at 1536 and swapping to disk.
  5. If it still OOMs, drop to a quant. Move to a Q8 GGUF build, launch ComfyUI with --lowvram, and make sure the machine has 32GB of system RAM. This is the route that gets you onto a 6GB card.
  6. Test with a fixed prompt and seed before you trust it. Generate the same prompt twice and confirm the output is stable. If two runs at the same seed differ substantially, you are on a workflow with a moving sampler setting rather than a hardware problem.

If you want extra control later, the community ecosystem is mature: an all-in-one ComfyUI pack collects the Turbo checkpoints and accessories under one Apache 2.0 collection, and Alibaba's own ControlNet Union model for the Z-Image Turbo Fun line adds structural conditioning for pose and depth.

Prompting an 8-Step Model

The prompting habits you built on 20–50 step models mostly transfer, with three adjustments.

Habit What to do with Z-Image Turbo
Long, stacked adjective prompts Write shorter. The model is trained for fast, direct prompts; name the subject, the setting, the light and the lens
Negative prompts Mostly unnecessary at CFG 1 — there is no guidance scale to push against. Drop it and see
30+ step sampling for detail Stay at 8 steps. More steps do not buy detail here; they just cost time
Generating at maximum resolution Generate at 1024 and upscale. On 8GB, off-grid or oversized canvases are where the failures live

Ready to try it yourself? Try Z-Image Free →

For subjects where you need product-grade output rather than a fast concept, the prompt structures in the Z-Image product photography guide are a good place to start — they were written for this family and they keep prompts short instead of stacking qualifiers.

How It Compares to Other Local Models

Z-Image Turbo FLUX.1 dev SDXL
Parameters 6B 12B ~3.5B (base + refiner)
Steps 8 20–50 20–30
Minimum practical VRAM ~8GB (FP8) / 6GB (GGUF) ~12GB (FP8), 24GB comfortable at full precision ~8GB comfortable
Licence Apache 2.0 Non-commercial Open, with per-model variations
Photorealism Very high Excellent Good, style-limited
Prompt adherence at low steps Strong Weak — needs steps Moderate

The comparison that matters for an 8GB card is not parameters, it is the licence and the step count. FLUX.1 dev is a strong model, but the non-commercial licence and the 24GB comfort zone put it out of reach for both the card and — for client work — the paperwork. Z-Image Turbo is the first model in my own experience where an 8GB card plus an Apache 2.0 licence is enough to run client work locally.

For a broader look at how this model stacks up against the current open-weight peers, the Z-Image vs Qwen Image comparison covers the quality trade-offs in more detail, and the Z-Image 2026 update overview tracks what changed after the base model release in January.

For a closer look at how it stacks up against other models, see [GLM](https://wanvideogenerator.com/blog/glm-image-vs-z-image?utm_source=blog&utm_medium=article&utm_campaign=z-image-turbo-fp8-local-guide).

Local or Hosted: The Honest Split

I run both, and the split is predictable.

Your situation Route Why
You have 8GB VRAM and 16GB+ RAM Local FP8 Free per image, offline, full control over seeds and workflows
You have a laptop with integrated graphics Hosted There is no quant small enough to make this pleasant
You need 100 product shots today Hosted Throughput beats per-image cost when you are on a deadline
You are iterating on a prompt Local 5–8 seconds per image makes exploration conversational
You need 2K output Generate at 1024 locally, then upscale Oversized local canvases are where OOM lives
You work on sensitive client images Local Nothing leaves your machine
You want anime or stylised output Z-Anime route Tuned variants exist and are cheaper than fighting the base model's photoreal bias
You need video, not images Hosted image to video Z-Image is a still-image family; video is a different toolchain

The pattern: local wins on iteration and privacy, hosted wins on volume, resolution and hardware-independence. Almost nobody needs only one of them.

Pros and Cons

Pros

  • Genuinely runs on 8GB VRAM in FP8 and 6GB with a GGUF quant.
  • Apache 2.0 licence — the cleanest route to commercial local generation at this quality.
  • 8 steps, no CFG: fast enough that prompt iteration feels interactive.
  • Mature ComfyUI support, an official ModelScope mirror and a healthy fine-tune ecosystem.
  • Strong photorealism for its parameter count.

Want to see the difference on your own footage? Start creating with Z-Image →

Cons

  • Not a reference-image model. Identity consistency across a set is a workflow problem, not an input slot.
  • Text-to-image only in the released checkpoints; editing variants are a separate branch of the family.
  • Quantisation still costs you fine texture, and Q4 is visibly softer.
  • AMD/ROCm performance is inconsistent and should be tested before you plan around it.
  • 1080p-class output is best reached by upscaling rather than native generation on small cards.

The Bottom Line

Skip the setup and test it in the browser: Experience Z-Image Free →

The interesting thing about Z-Image Turbo is not the benchmark score. It is that the barrier moved. An 8GB card with 16GB of system RAM, a permissive licence and an 8-step sampler is now enough to run real image work locally, and that was not true a year ago.

Install FP8, keep it at 1024, stay at 8 steps, and use a quant only if your card demands one. Then decide per job whether the render belongs on your machine or on a hosted one — because the correct answer to "local or cloud" is almost always "whichever one is faster for today's task."

FAQ

How much VRAM does Z-Image Turbo need? About 14–16GB for the BF16 build, roughly 8GB for the FP8 build, and around 6GB for community GGUF quants with --lowvram. System RAM matters as well: 16GB is the practical minimum for FP8 and 32GB is recommended for the GGUF low-VRAM route.

Does FP8 reduce image quality? Slightly, and less than most people expect. In side-by-side tests at 1024px, FP8 and BF16 were effectively indistinguishable in composition and prompt adherence, with a small softening in fine texture. Q4 quants are where the difference becomes obvious.

Is Z-Image Turbo free to use commercially? It is released under Apache 2.0, which is a permissive licence that generally allows commercial use. Read the licence file in the model repository yourself before shipping client work — this guide is a summary, not legal advice.

Can I run it in ComfyUI on an 8GB card? Yes, with the FP8 checkpoint. Place the checkpoint in models/diffusion_models, the text encoder and VAE in their own folders, and use the Z-Image template workflow at 8 steps with CFG 1. Ensure the machine has at least 16GB of system RAM.

Why is my 8GB card still running out of memory? Usually because BF16 weights are loaded, the resolution is too large, or system RAM is the real constraint. Switch to FP8, drop to 1024px, and check RAM usage. If it still fails, move to a Q8 GGUF build and launch with --lowvram.

Does raising the step count improve the output? No, and it costs you the model's main advantage. Z-Image Turbo was trained for the 8-step, no-CFG regime; additional steps add time without adding detail.

Can Z-Image Turbo keep a character consistent across images? Not through a reference-image input — the released checkpoints are text-to-image. Consistency comes from prompt locking, pairing Turbo with an edit model, or training a LoRA (which the base model is designed to support).

Which is better for client work, local or hosted? Local for iteration, privacy and zero marginal cost; hosted for volume, higher resolutions and machines without a capable GPU. Most people doing real work end up using both, and picking per job.

References

Related guides

Start Generating

Ready to Generate Images with Z-Image?Generate with Z-Image

Use Z-Image to create images, edits and variations — start free in your browser.

Text to Image
Image to Image
Free to Try
No Setup Required