One bet, one set of prompts, one set of seeds, one frozen instrument, one statistic, one set of thresholds. Run on the first generator, naming a position controlled where the subject went. Run on the second, it did not. This paper is about what that difference is, and is not. It is not that one engine understands spatial language and the other does not: both move the subject the right way when told. The difference is whether the movement is repeatable enough to be called control, and the line between the two engines falls exactly where signal stops exceeding noise.
- What was held identical, and what was not
- The six cells, twice
- The words get through on both engines
- Signal against noise: where the engines part
- The variance changed hands, it did not shrink
- What both engines share
- What genuinely differs
- Threats to this comparison
- What this means for reading prompt studies
- Files and reproduction
- Appendix
1. What was held identical, and what was not
The second engine was run as a replication, not as a new study: the plan, the prompts, the seeds, the instrument and the decision rules were fixed before the first image and inherited without edit. The intent was a single-variable comparison. It is not quite one, and the table says exactly where it falls short.
| Element | Z-Image Turbo | FLUX.1-schnell | |
|---|---|---|---|
| Prompts | 27 position-named + 27 plain, identical text | same | held |
| Seeds | 50000, 154729, 259458, 364187, 468916, 573645 | same | held |
| Resolution | 768 × 768 | same | held |
| Steps / cfg | 4 / 1.0 | same | held |
| Instrument | objmask.py, sha256-frozen | same | held |
| Statistic and thresholds | subject-centred ICC(A,1), 95% McGraw-Wong | same | held |
| Referee rules | R1–R6 as declared | same | held |
| Model | z_image_turbo_bf16.safetensors | flux1-schnell-Q5_K_S.gguf | differs |
| Weight precision | bf16 (full) | 5-bit quantised | differs |
| Text encoder | Qwen3-4B | T5-XXL at 4-bit + CLIP-L | differs |
| Sampler / schedule | res_multistep, shift 3 | euler, simple | differs |
| Median render | 43.6 s | 82.8 s | differs |
The corpora differ in one further way. The original design had a third prompt family naming size; on the second engine a budget rule fixed in advance dropped it, because the pilot timing put the full set past the declared limit. Six cells are therefore common to both engines and this paper compares only those. The comparison rules themselves, including what counts as a replication, were written and committed before any image from the second engine was measured.
2. The six cells, twice
| Family | Axis | Expected | Z-Image P | FLUX P | ΔP | Verdict | ||
|---|---|---|---|---|---|---|---|---|
| position named | horizontal position | reach | 0.838 | confirms | 0.071 | FAILS | -0.767 | DIFFERS |
| position named | vertical position | reach | 0.497 | open | 0.182 | FAILS | -0.315 | underpowered |
| position named | size | not | 0.634 | open | 0.072 | confirms | -0.562 | DIFFERS |
| plain prompts | horizontal position | not | 0.064 | confirms | 0.000 | confirms | -0.064 | replicates |
| plain prompts | vertical position | not | 0.169 | confirms | 0.115 | confirms | -0.054 | replicates |
| plain prompts | size | not | 0.506 | open | 0.106 | confirms | -0.400 | DIFFERS |
2 of 6 cells replicate, and both are cells where neither engine moves
the subject, which is agreement about an absence. The pre-registered headline, that horizontal position is
reached and the plain-prompt axes are not, holds on the first engine and not on the second
(headline_holds: false in the committed comparison file). The
horizontal position cell moves from 0.838 to
0.071, a drop of 0.767, and crosses from
confirming the prediction to failing it.
Read on its own, that table invites a conclusion this paper is going to reject: that the second engine does not understand where “top-left” is.
3. The words get through on both engines
The variance share asks a demanding question: does a prompt put the subject in the same place every time. A blunter question is whether the instruction moves the average at all, and in the right direction. Regress the measured position on the instructed level, scoring left and top as −1, centre as 0, right and bottom as +1. An engine placing subjects exactly in the centre of each instructed third would score about 0.33; an engine ignoring the words entirely would score 0.
| Engine | Axis | Slope | told left / top | told centre / middle | told right / bottom |
|---|---|---|---|---|---|
| Z-Image Turbo | horizontal position | +0.157 | -0.136 | +0.010 | +0.178 |
| FLUX.1-schnell | horizontal position | +0.066 | -0.033 | +0.001 | +0.098 |
| Z-Image Turbo | vertical position | +0.131 | -0.037 | +0.108 | +0.225 |
| FLUX.1-schnell | vertical position | +0.094 | +0.024 | +0.147 | +0.212 |
Every slope is positive. On the second engine the instruction moves the mean by +0.066 horizontally and +0.094 vertically, against +0.157 and +0.131 on the first. Vertically the second engine retains 72% of the first engine's response; horizontally, 42%.
There is a subtlety worth stating, because it cuts against a claim that is easy to make from hit rates alone. On the vertical axis the second engine's upward response is the larger of its two: told “top” it lifts the subject 0.123 from where “middle” puts it, and told “bottom” it lowers it 0.064. It still never lands in the upper third, because its default composition sits low (+0.147) and the shift is not large enough to clear the boundary. A real response scores as a miss. The bias is in where the engine starts, not in whether it listens.
4. Signal against noise: where the engines part
If both engines move the subject the right way, the failure must be in consistency. Split the spread of each measurement into the part that separates one prompt from another, and the part that is left once both prompt and seed are accounted for: the scatter of the same prompt on different seeds.
| Engine | Family | Axis | Prompt SD | Residual SD | Signal ÷ noise |
|---|---|---|---|---|---|
| Z-Image Turbo | position named | horizontal position | 0.124 | 0.051 | 2.42 |
| FLUX.1-schnell | position named | horizontal position | 0.056 | 0.106 | 0.53 |
| Z-Image Turbo | position named | vertical position | 0.117 | 0.099 | 1.18 |
| FLUX.1-schnell | position named | vertical position | 0.098 | 0.144 | 0.68 |
| Z-Image Turbo | position named | size | 0.117 | 0.071 | 1.64 |
| FLUX.1-schnell | position named | size | 0.064 | 0.114 | 0.56 |
| Z-Image Turbo | plain prompts | horizontal position | 0.011 | 0.021 | 0.52 |
| FLUX.1-schnell | plain prompts | horizontal position | 0.014 | 0.037 | 0.37 |
| Z-Image Turbo | plain prompts | vertical position | 0.023 | 0.034 | 0.66 |
| FLUX.1-schnell | plain prompts | vertical position | 0.039 | 0.064 | 0.61 |
| Z-Image Turbo | plain prompts | size | 0.055 | 0.040 | 1.36 |
| FLUX.1-schnell | plain prompts | size | 0.053 | 0.073 | 0.72 |
This is the cleanest statement of the difference the study found. On the position-named family the first engine's horizontal signal is 2.42 times its noise and its vertical signal 1.18 times; the second engine scores 0.53 and 0.68. One engine is above parity on both axes, the other below on both. Everything the verdicts said follows from that crossing.
The residual is the operative term. It doubles horizontally, from 0.051 to 0.106, and rises by about half vertically, 0.099 to 0.144. The same prompt on six different seeds simply lands in six more different places. Prompt control does not fail here because the instruction is weak; it fails because the instruction is drowned.
That figure is the result. Nothing in the lower row is a refusal to follow the instruction, and several of its frames do sit the vase left of centre. What is missing is agreement between them.
5. The variance changed hands, it did not shrink
A collapsed prompt share is easy to picture as a subject that stopped moving. The totals say otherwise.
| Engine | Family | Axis | Total spread (SD) | Prompt share | Seed share | Residual share |
|---|---|---|---|---|---|---|
| Z-Image Turbo | position named | horizontal position | 0.134 | 0.85 | 0.00 | 0.15 |
| FLUX.1-schnell | position named | horizontal position | 0.135 | 0.18 | 0.21 | 0.61 |
| Z-Image Turbo | position named | vertical position | 0.159 | 0.54 | 0.06 | 0.39 |
| FLUX.1-schnell | position named | vertical position | 0.187 | 0.27 | 0.14 | 0.59 |
| Z-Image Turbo | position named | size | 0.144 | 0.66 | 0.09 | 0.25 |
| FLUX.1-schnell | position named | size | 0.161 | 0.16 | 0.34 | 0.51 |
| Z-Image Turbo | plain prompts | horizontal position | 0.025 | 0.19 | 0.12 | 0.69 |
| FLUX.1-schnell | plain prompts | horizontal position | 0.053 | 0.07 | 0.45 | 0.49 |
| Z-Image Turbo | plain prompts | vertical position | 0.044 | 0.27 | 0.12 | 0.61 |
| FLUX.1-schnell | plain prompts | vertical position | 0.086 | 0.21 | 0.23 | 0.56 |
| Z-Image Turbo | plain prompts | size | 0.074 | 0.55 | 0.15 | 0.30 |
| FLUX.1-schnell | plain prompts | size | 0.130 | 0.16 | 0.52 | 0.31 |
Total spread for horizontal position under position-named prompts is 0.134 on the first engine and 0.135 on the second: the same, to three decimals. Vertically the second engine actually spreads wider (0.187 against 0.159). Subjects move as much; what changes is who decides where they go.
Under position-named prompts the first engine gives the prompt 85% of horizontal position and leaves 15% to the residual. The second engine reverses it: 18% to the prompt, 61% to the residual. The largest single term on the second engine is not the prompt and not the seed, but the particular pairing of the two, which by construction neither party can be said to have chosen.
6. What both engines share
The differences are easier to see than the agreements, and the agreements are the more useful half, because a property that survives an engine change is a candidate for a property of this class of model.
| Engine | Default vertical position | Hit rate, centre / bottom-centre | Size–displacement r, named / plain | Subjects in central column, plain prompts |
|---|---|---|---|---|
| Z-Image Turbo | +0.088 | 94% / 100% | -0.42 / -0.18 | 100% |
| FLUX.1-schnell | +0.128 | 56% / 72% | -0.65 / -0.32 | 99% |
Only the centre column works, on either engine
This is the finding most likely to be misread from the first engine alone. Its headline horizontal share of 0.84 does not mean it puts subjects where it is told: it means each prompt moves the subject to the same place every time, and off centre that place is frequently the wrong one. Measured as obedience rather than repeatability, the first engine scores 64% horizontally and the second 51%, against 33% chance. Only “centre” and “bottom-centre” work reliably, on either engine.
Both compose low
With nothing named, both engines place the subject slightly below the middle of the frame (+0.088 and +0.128) and near the horizontal centre. Neither lifts a subject into the upper third on request, though as section 3 showed, both do shift upward when asked.
Displacement requires shrinkage, on both
Hold a prompt fixed and vary only the seed: on both engines, the renders that pushed the subject furthest from the centre are the renders that made it smallest (r = -0.42 and -0.65 with a position named, -0.18 and -0.32 with nothing named). It survives in plain prompts, where no wording mentions position or size at all, so it is a property of how these generators compose a frame rather than a response to instructions.
7. What genuinely differs
| Engine | Horizontal obedience | By subject: vase / cat / man | Seed share, plain prompts | Prompts dropped | Audit masks wrong |
|---|---|---|---|---|---|
| Z-Image Turbo | 64% | 80% / 49% / 63% | 0.12 | 3 of 27 | 5 of 30 |
| FLUX.1-schnell | 51% | 51% / 51% / 51% | 0.45 | 6 of 27 | 6 of 30 |
Subject dependence disappears. On the first engine what you are placing matters: a rigid vase obeys far better than a cat or a person. On the second the three subjects are within a point of each other. That flatness is itself the signature of an instruction that is not being acted on; when placement is working, the affordances of the subject show up in the numbers.
The seed matters more, but only where nothing is named. Under plain prompts the seed share rises from 0.12 to 0.45 horizontally. Read with the note in section 5: it is a bigger share of a smaller spread.
The instrument struggled more on the second engine, and asymmetrically: 6 position-named prompts were dropped whole against 3, and the visual audit counted 6 wrong masks against 5, at a declared limit of 6. Section 8 treats this as a threat rather than a finding.
And they cost differently. At matched resolution and step count the first engine renders in a median 43.6 s against the second's 82.8 s on the same machine, so the engine with the working position control is also roughly 1.9 times faster here. That is a property of these builds on this hardware, not of the architectures, but it is the trade a person actually faces: the second engine's weights are quantised to fit a 32 GB machine, and the fitting is part of what is being measured.
8. Threats to this comparison
- Four variables move together. Model, weight precision, text encoder and sampler all differ. Section 3 rules out the encoder hypothesis in its strongest form, since the words demonstrably arrive, but it cannot apportion the rest. The clean follow-up is to raise the second engine's encoder precision and change nothing else.
- Each engine ran at its own recommended settings. Matching samplers would isolate the model better and would also mean running at least one engine in a configuration nobody uses.
- The instrument failed asymmetrically. Every wrong mask in the second engine's audit fell in the family that failed, and its audit passed at exactly the declared limit. The second engine's own paper tests this three ways and finds the failure survives, but the precise split between prompt, seed and residual should be read as approximate on that engine.
- Drop rates differ. The whole-prompt rule removed 36 images from the second engine's position family and 18 from the first, and on the second engine the losses skew toward off-centre instructions. Neither family is complete, and they are incomplete in different places.
- Six cells, not nine. Nothing here compares how the two engines handle a size instruction, because that family was never generated on the second.
- One operating point each, three subjects, one framing vocabulary. All prompts are studio-style single-subject descriptions on plain backgrounds.
9. What this means for reading prompt studies
The useful claim from this pair of studies is not about either engine. It is that prompt steerability of geometry is a property of the configuration you are running, and it is not safe to generalise from one. The same sentence, the same seeds, the same measurement and the same thresholds produced a clean control on one locally-runnable setup and a failed one on another, at matched resolution, steps and guidance. A paper reporting that spatial words control composition in diffusion models, on the evidence of one model, is reporting that model.
Three narrower things travel further than that:
- Report signal against noise, not only a share. A variance share compresses two different failures, a weak instruction and an inconsistent one, into a single number. The prompt-to-residual ratio separates them, and here it is what actually distinguishes the engines.
- Do not read a high share as obedience. The first engine's 0.84 came with 64% accuracy. It is repeatable, not correct. Both engines have a usable window of two instructions out of nine.
- Named axes are not independent. Position and size are coupled through frame occupancy on both engines. Any design that assumes a prompt can name one without moving the other is mis-specified, and this bet was.
The bet that produced all of this was killed on the second engine. It was worth writing precisely because it could be: the frozen thresholds, the carried-over instrument and the committed comparison rules are what make a reversal legible as a reversal rather than as a change of mind.
10. Files and reproduction
This paper reads both studies' result files and retypes no number. Each study reproduces from its own folder, and each has its own paper covering it alone.
| Item | Where |
|---|---|
| First engine: study, verdict, paper | named-axes/, zimage-paper/paper.html |
| Second engine: study, verdict, paper | named-axes-flux/, flux-paper/paper.html |
| Comparison rules, committed before the second engine was measured | named-axes-flux/compare_engines.py, comparison.json |
| This paper | compare-paper/build_paper.py |
Corpus: 486 images on the first engine and 324 on the second, 810 in total, every one measured once by the same frozen instrument.
Appendix
A1. The two configurations in full
| Element | Z-Image Turbo | FLUX.1-schnell | |
|---|---|---|---|
| Prompts | 27 position-named + 27 plain, identical text | same | held |
| Seeds | 50000, 154729, 259458, 364187, 468916, 573645 | same | held |
| Resolution | 768 × 768 | same | held |
| Steps / cfg | 4 / 1.0 | same | held |
| Instrument | objmask.py, sha256-frozen | same | held |
| Statistic and thresholds | subject-centred ICC(A,1), 95% McGraw-Wong | same | held |
| Referee rules | R1–R6 as declared | same | held |
| Model | z_image_turbo_bf16.safetensors | flux1-schnell-Q5_K_S.gguf | differs |
| Weight precision | bf16 (full) | 5-bit quantised | differs |
| Text encoder | Qwen3-4B | T5-XXL at 4-bit + CLIP-L | differs |
| Sampler / schedule | res_multistep, shift 3 | euler, simple | differs |
| Median render | 43.6 s | 82.8 s | differs |
Both ran in headless ComfyUI 0.17.2 on an Apple M1 Max with 32 GB, writing PNGs at 768 × 768. Negative conditioning is inert on both, since cfg is 1. Seeds are the same six integers in the same order, and every prompt in a family was rendered on all six.
A2. The shared prompt set
Two families of 27, each three subjects crossed with nine variants, identical text on both engines. The subjects are a blue ceramic vase, a black cat, and a man in a red jacket: one rigid object, one deformable animal, one person. One example from each family:
| Family | First prompt |
|---|---|
| position named | A photograph of a blue ceramic vase in the top-left corner of the frame, with a plain background. |
| plain prompts | A photograph of a blue ceramic vase, with a plain background. |
The nine position phrases are top-left, top-centre, top-right, mid-left, centre, mid-right, bot-left, bot-centre, bot-right. The nine plain variants are neutral paraphrases that change the wording without naming any geometry.
A3. The two ratios this paper uses
Prompt share (P) is an intraclass correlation, ICC(A,1), in a crossed prompt-by-seed design with each subject's mean removed: of the spread in one measurement, the fraction attributable to which prompt was used, reported with a 95% McGraw and Wong interval. It is the statistic the bet was written on, and the only one that decides a verdict.
Signal ÷ noise is the prompt standard deviation divided by the residual standard deviation from the same decomposition. It is not a verdict statistic and appears nowhere in either bet; it is used here because a share compresses two distinct failures, a weak instruction and an inconsistent one, into one number, and the comparison turns on telling them apart. Both are computed from the same subject-centred matrices.
A note on reading the descriptive shares in section 5 next to the verdict shares in section 2: the descriptive split credits sampling noise in a prompt's mean to the prompt, so it reads higher than the corrected ICC. The ordering of the three terms is unaffected.
A4. Referee checks, both engines
| Engine | R3 determinism | R4 seed crossing | R5 validity (position family) | R6 visual audit | Verdict |
|---|---|---|---|---|---|
| Z-Image Turbo | pixel-identical | pass | 2.47% invalid | 5 of 30 wrong | INCONCLUSIVE |
| FLUX.1-schnell | pixel-identical | pass | 4.32% invalid | 6 of 30 wrong | KILLED (failure) |
R1 (instrument on synthetic images) and R2 (interval coverage) test the instrument and the statistic rather than the engine, so the second study carried over the first study's results rather than re-running them, and recorded them as carried over. Every check passed on both engines, which is what makes the differing verdicts attributable to the engines.
A5. Terms
| Term | Meaning here |
|---|---|
| Cell | one prompt family crossed with one measured axis; six are common to both engines |
| Reach / leak | a named axis the prompt is predicted to control / an unnamed axis it controls anyway |
| Replicates | both engines reached the same status and their intervals overlap; fixed before measurement |
| Hit rate, obedience | fraction of images where the subject occupies the instructed third; unrelated to P |
| Slope | change in mean measured position per step of instructed position; a directional test |
| Residual | what is left after prompt and seed: the pairing of the two, plus measurement error |
A6. Reproduction
Each study rebuilds from its own folder; this paper rebuilds from both. In order, per engine:
generate · rerender · r3 · measure
· audit · analyze, then
python3 compare_engines.py for the committed comparison, then
python3 compare-paper/build_paper.py for this document. The instrument's sha256 is checked
against a freeze on every run and the analysis refuses to proceed if it differs.