Z-Image#
Z-Image is a Lumina2-architecture family of seven models built around long, natural-language prompts. Pick one when you want a modern DiT look, very fast turbo variants, or unlimited LoRA stacking.
Three of these live under the Z Image family entry in the model picker, three under Z Anime, and ZiT-ANI is a direct pick.
At a glance#
| Model | Picker group | Steps | CFG | Sampler | Scheduler | Base size | Pick it for |
|---|---|---|---|---|---|---|---|
| Z Image Turbo | Z Image | 12 | 1 | euler |
simple |
1024x1536 | The everyday Z-Image. Fast, general purpose. |
| Cyber Realistic Turbo | Z Image | 12 | 1 | dpmpp_2m |
beta |
1280x1920 | Photoreal look on the largest native canvas in the family. |
| Z Image Base | Z Image | 35 | 4 | res_multistep |
simple |
1024x1536 | Prompt obedience and a CFG dial that actually does something. |
| Z Anime Base | Z Anime | 35 | 4 | euler_ancestral |
beta |
1024x1536 | Painterly anime at full quality. Widest CFG range here (3-9). |
| Z Anime 8-Step | Z Anime | 8 | 1 | euler_ancestral |
beta |
1024x1536 | The same painterly look for a third of the wait. |
| Z Anime 4-Step | Z Anime | 4 | 1 | euler_ancestral |
beta |
1024x1536 | The fewest steps in the family, when you want throughput. |
| ZiT-ANI | direct pick | 12 | 1 | euler |
simple |
1024x1536 | An anime-leaning alternative to Z Image Turbo on the same recipe. |
All seven default to the Remacri (4x) upscaler and a batch of 2, and none of them add quality tags to your prompt.
Steps and CFG are clamped, not rejected
If you pass a value outside a model's range, you get an image at the nearest legal value rather than an error. See Parameters.
Prompt style#
Z-Image wants one flowing paragraph of natural language, roughly 150 to 250 words. The text encoder truncates beyond that, so longer prompts lose their tail. Very short prompts (under about 25 words) leave the model to invent most of the frame.
None of the Z-Image models add quality tags for you. Whatever you type is what gets encoded, so masterpiece, best quality, 8k is wasted tokens here.
Rules that matter in practice:
| Do | Don't |
|---|---|
| Write sentences, not comma-separated booru tags | Use 1girl, solo, looking_at_viewer style tag lists |
| Phrase constraints positively: "sharp focus, crisp details" | Write "no blur", or rely on the negative prompt |
| Wrap any text you want rendered in double quotes | Use weight syntax like (word:1.5) -- it does nothing |
| Lead with the camera and composition, then subject, then lighting | Add meta tags: 8K, masterpiece, trending on artstation |
Character names are rewritten into natural form on this family -- Hatsune Miku, not hatsune_miku \(vocaloid\). See Character matching.
If you turn on AI enhancement, Eimi uses a Z-Image-specific instruction set that produces exactly this shape and caps the result at one paragraph. Z-Anime models get their own variant. See AI Enhance and the prompting guide.
Negative prompts do not apply#
Negative prompts have no useful effect here
Eimi adds nothing of its own to the negative prompt on Z-Image models, and the family runs at CFG 1 on every turbo variant, where a negative has effectively no influence. Anything you type in negative: is passed through verbatim and then ignored by the model. The only exception is Z Anime Base, which ships a short built-in negative (lowres, bad quality, worst quality, jpeg artifacts, ugly, blurry, watermark, signature).
The two 35-step models (Z Image Base and Z Anime Base) run at CFG 4 by default, so a negative does have some pull on them. On the five CFG-1 models it does not.
Express what you want, not what you don't.
The seven models#
Z Image Turbo#
The distilled 12-step build and the one to start with. If you want a Z-Image render and have no particular reason to pick another, pick this.
| Architecture | Lumina2 (Z-Image) |
| Subtype | none |
| Prompting | Prose, one paragraph, roughly 150-250 words |
| Negatives | Passed through verbatim, inert at CFG 1 |
| Steps | 12 (6-15) |
| CFG | 1 (1-2) |
| Sampler | euler |
| Scheduler | simple |
| Base resolution | 1024x1536 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: natural-language scene descriptions, fast turnaround at 12 steps, the widest LoRA support in the bot (44 of the 91 LoRAs list Z-Image), the full 11-ratio resolution table including 21:9 and 9:21.
Weak at: anything you want to exclude -- the negative prompt is dead weight. Being a distilled model, it also gives you noticeably less composition variety from seed to seed than an undistilled model would; re-rolling the same prompt tends to hand you the same picture with different details. The CFG slider is nearly inert across its 1-2 range. And because the whole family refuses the Detailer, a render with mangled hands or a soft background face has to be re-rolled rather than repaired.
Cyber Realistic Turbo#
A photoreal finetune on the standard 12-step turbo recipe, running on the biggest native canvas in the family.
| Architecture | Lumina2 (Z-Image) |
| Subtype | none |
| Prompting | Prose, one paragraph, roughly 150-250 words |
| Negatives | Passed through verbatim, inert at CFG 1 |
| Steps | 12 (6-15) |
| CFG | 1 (1-3) |
| Sampler | dpmpp_2m |
| Scheduler | beta |
| Base resolution | 1280x1920 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: photographic subjects, skin and fabric texture, and delivering a large image straight out of the sampler without a hi-res pass. Its 21:9 entry is 2400x1024, the widest native frame in the family.
Weak at: anime and flat illustration -- ask it for cel shading and you get a photo of a cosplayer. It is the most expensive model in the family: about 1.6x the pixels of the other six at the same aspect ratio, which costs VRAM, render time, and upload time on the way back to Discord. Batch 2 at the default 1280x1920 is a meaningful chunk of one GPU. Negatives are inert here as on every other turbo variant, and the Detailer refusal bites hardest on this model because face detail at 1280x1920 is exactly what you would want to repair.
Z Image Base#
The undistilled checkpoint. Slower, more literal, and the only non-anime model here where CFG behaves like a real dial.
| Architecture | Lumina2 (Z-Image) |
| Subtype | none |
| Prompting | Prose, one paragraph, roughly 150-250 words |
| Negatives | Passed through verbatim, some effect at CFG 4 |
| Steps | 35 (25-55) |
| CFG | 4 (3-5) |
| Sampler | res_multistep |
| Scheduler | simple |
| Base resolution | 1024x1536 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: following a long prompt clause by clause, unusual compositions the turbo build flattens out, and giving genuinely different results across seeds.
Weak at: speed. At 35 steps it is close to three times the sampling work of Z Image Turbo for the same picture size, and it lifts the GPU power cap while it runs, so the machine gets louder and hotter than it does for the rest of the family. It also has no baked-in style, so plain prompts come back plainer than the finetunes give you -- you have to describe the look you want. Detailer and Ultimate SD Upscale still refuse it.
Z Anime Base#
The painterly anime finetune at full quality. Soft brushwork and illustrated backgrounds rather than flat lineart.
| Architecture | Lumina2 (Z-Image) |
| Subtype | z-anime |
| Prompting | Art-direction prose, roughly 60-280 words |
| Negatives | Ships its own short list; yours is honoured and has some effect at CFG 4 |
| Steps | 35 (28-50) |
| CFG | 4 (3-9) |
| Sampler | euler_ancestral |
| Scheduler | beta |
| Base resolution | 1024x1536 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: painted anime illustration, atmospheric backgrounds, and being steered -- the 3 to 9 CFG range is the widest in the family, so you can push prompt adherence far past what any other Z-Image model allows. It is the only model here that ships a built-in negative prompt.
Weak at: flat cel-shaded, line-art anime -- that is SDXL territory, and Z Anime will soften it every time. It is slow at 35 steps and lifts the GPU power cap while it runs. Pushing CFG toward 9 buys adherence at the cost of burnt, over-saturated colour. No Detailer, no USDU.
Z Anime 8-Step#
Z Anime Base distilled to 8 steps at CFG 1. Same look, roughly a quarter of the sampling work.
| Architecture | Lumina2 (Z-Image) |
| Subtype | z-anime |
| Prompting | Art-direction prose, roughly 60-280 words |
| Negatives | Passed through verbatim, inert at CFG 1 |
| Steps | 8 (8-10) |
| CFG | 1 (1-1.5) |
| Sampler | euler_ancestral |
| Scheduler | beta |
| Base resolution | 1024x1536 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: the painterly Z-Anime look without the 35-step wait. This is the sensible default of the three Z Anime entries.
Weak at: everything distillation costs. The 3-9 CFG range collapses to 1-1.5, the built-in negative is gone, and seed-to-seed variety narrows sharply -- re-rolling gives you variations rather than alternatives. Fine detail at the edge of the frame is softer than the base model's.
Note
This build clears VRAM before it loads, so the first job after switching to it can sit waiting a little longer than usual. Subsequent jobs on the same model do not pay that cost.
Z Anime 4-Step#
The fewest sampling steps of any model in this family. Four steps, CFG 1.
| Architecture | Lumina2 (Z-Image) |
| Subtype | z-anime |
| Prompting | Art-direction prose, roughly 60-280 words |
| Negatives | Passed through verbatim, inert at CFG 1 |
| Steps | 4 (4-6) |
| CFG | 1 (1-1.5) |
| Sampler | euler_ancestral |
| Scheduler | beta |
| Base resolution | 1024x1536 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: iterating. When you are hunting for a composition and want ten tries in the time one Z Anime Base render takes, this is the model.
Weak at: finishing. Four steps is not enough sampling for hands, small faces, crowd scenes, rendered text or intricate patterns, and all of those come back visibly broken more often than not. Seed variety is the narrowest in the family. And with the Detailer refusing the whole family, there is no repair pass available for exactly the failures this model produces most. Treat its output as a sketch, then re-render the prompt you liked on Z Anime 8-Step or Base.
ZiT-ANI#
An anime-leaning Lumina2 checkpoint on the standard 12-step turbo recipe. A direct pick, not part of either family group.
| Architecture | Lumina2 (Z-Image) |
| Subtype | lumina2 |
| Prompting | Prose, one paragraph, roughly 150-250 words |
| Negatives | Passed through verbatim, inert at CFG 1 |
| Steps | 12 (6-15) |
| CFG | 1 (1-2) |
| Sampler | euler |
| Scheduler | simple |
| Base resolution | 1024x1536 |
| Default batch | 2 |
| Default upscaler | Remacri (4x) |
Good at: an anime slant on the Z Image Turbo speed profile, without the painterly softness the Z Anime models impose.
Weak at: the same things Z Image Turbo is weak at, since the two run an identical recipe -- 12 steps, CFG 1, euler/simple, 1024x1536, Remacri. The only difference is the checkpoint's training, so the choice between them is a taste test, not a settings decision. It is also the thinnest-documented model in the family: its config entry says nothing beyond the sampler recipe, so there is no author guidance to pass on. Distilled behaviour, dead negatives, no Detailer, no USDU.
Unlimited LoRAs#
Z-Image is the only family with no LoRA slot limit. Every other family caps you at 4 LoRAs per job; Z-Image loads as many as you select, affecting both the model and the text encoder.
It is also the best-supported family in the LoRA library -- 44 of the 91 LoRAs list Z-Image compatibility, more than any other architecture. See the LoRA catalog.
The pixel LoRA overrides your resolution
Selecting the pixel_6x6 LoRA forces the render to 768x768 (1:1) regardless of what you asked for, and adds two pixel-quantization passes (6x6 grid, 48 colors) after decode. That is the intended behavior, but it silently discards your aspect ratio.
Resolutions#
Six of the seven models share one table. The default is 1024x1536 (2:3).
| 1:1 | 9:7 | 7:9 | 4:3 | 3:4 | 3:2 | 2:3 | 16:9 | 9:16 | 21:9 | 9:21 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1280x1280 | 1440x1120 | 1120x1440 | 1472x1104 | 1104x1472 | 1536x1024 | 1024x1536 | 1600x896 | 896x1600 | 1680x720 | 720x1680 |
Cyber Realistic Turbo renders larger, defaulting to 1280x1920:
| 1:1 | 9:7 | 7:9 | 4:3 | 3:4 | 3:2 | 2:3 | 16:9 | 9:16 | 21:9 | 9:21 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1568x1568 | 1776x1376 | 1376x1776 | 1808x1360 | 1360x1808 | 1920x1280 | 1280x1920 | 2096x1168 | 1168x2096 | 2400x1024 | 1024x2400 |
You can also type a literal WIDTHxHEIGHT. It is snapped down to a multiple of 8, floored at 64, and scaled down proportionally if the longest side goes past 4096.
Z-Anime#
The three Z-Anime models are painterly anime finetunes. They do not produce the flat, line-art-and-cel look that the SDXL family gives you; expect soft brushwork, blended shading and illustrated backgrounds.
Their prompt style is art direction prose of roughly 60 to 280 words. Booru tags, 8K / 4K / trending on artstation, and camera and lens specs are all counterproductive unless you specifically want a photo look.
When AI enhancement is on, the enhancer picks between four genre templates based on what you asked for:
| Template | Shape of the output |
|---|---|
| Character portrait | Detailed anime portrait of the character, soft rim lighting, expressive eyes with detailed reflections, fine hair strands, clean linework. |
| Action scene | Dynamic scene, dramatic angle, motion energy, speed lines, particle effects, cinematic composition. |
| Background / landscape | Location at a stated time of day, named lighting and atmosphere, Studio Ghibli level of background detail, wallpaper quality. |
| Full scene with characters | Opens with the subject and their traits, then the pose or action, then the setting. |
For a known character, the enhancer also adds a canon-anchor block describing the character's fixed features so the finetune does not drift.
What does not work on Z-Image#
Detailer and Ultimate SD Upscale both refuse to run
Both are blocked for the whole Z-Image family and you get an error message instead of a job:
- The Detailer (face, eye, hand and person refinement passes) rejects Z-Image models outright.
- Ultimate SD Upscale (the tiled diffusion upscale method) rejects them too. Use the model-based upscale or the refine method instead.
The Detailer refusal is the one that costs you something real. On every family that accepts it, a soft face or a six-fingered hand is a one-click repair. Here it is a re-roll, which is why the very fast models -- 4-Step especially -- are better used to find a composition than to finish one.
Everything else works normally: hi-res fix, model-based /upscale, SeedVR2, remix, and the /turbo MrFlow presets built on Z Image Turbo.
Related#
- Model overview -- the full catalog and how the picker is organised
- Upscaler catalog -- Remacri is this family's default upscaler
- /imagine -- every parameter you can pass
- /turbo -- three MrFlow presets run on Z Image Turbo