When the words stop steering

A pre-registered bet about prompt control, killed on FLUX.1-schnell

324 generated images, one measurement pass, six referee checks, all passed

Russell Parrish / Parallax Metrology · generated and analysed 2026-09-15 to 2026-09-16 · internal research paper, not for circulation

A bet was written down before any image existed: within a subject, a prompt controls exactly the geometric axes it names, and leaves the others to the sampler's draw. On the engine it was written for it came through without a single cell failing. Run unchanged on FLUX.1-schnell, it died. Position words do not place the subject here. This paper reports what happened instead, and what the failure exposes about the assumption underneath the bet: that position and size are separate axes at all.

Contents
  1. Setup: engine, corpus, instrument, protocol
  2. How the run was executed
  3. The verdict
  4. The operating window
  5. Repeatability is not obedience
  6. Who owns the variance
  7. Displacement requires shrinkage
  8. The default composition
  9. Is it the engine, or the instrument?
  10. What this suggests testing next
  11. Limits, and what this does not prove
  12. Files and reproduction
  13. Appendix: the recipe

1. Setup

The engine and the corpus

Every image comes from FLUX.1-schnell, quantised to Q5_K_S GGUF (8.26 GB, city96 build, Apache-2.0), with a 4-bit T5-XXL text encoder and CLIP-L, running locally in headless ComfyUI 0.17.2 on an Apple M1 Max with 32 GB. Images are 768 × 768, 4 steps, cfg 1, euler with the simple schedule. Generation is deterministic: three re-rendered cells came back pixel-identical to their originals (R3, section 3).

SetPromptsSeedsImagesPurpose
Position named27 (3 subjects × 9 instructed positions)6162does a position word place the subject?
Plain prompts27 (9 neutral paraphrases × 3 subjects)6162the engine's default composition

The same six seeds run across every prompt, so a seed's own effect is separable from the prompt's. The three subjects are a blue ceramic vase, a black cat, and a man in a red jacket: one rigid object, one deformable animal, one person.

A third family was dropped before any image was generated. The original design had a size-named family as well. A pilot measured 93.3 s per image on this engine, which put all three families at 11.5 hours against a budget rule of 8 hours fixed in advance. The rule named the fallback in advance too: keep position-named and plain, drop size-named. So FLUX carries six cells rather than nine, and the full-study “survives” label was unreachable here by construction. The failure reported below is in the position-named family, which was generated in full.

The instrument

One reader measures every image: an object-level mask that finds the subject against a background model fitted to the frame border, returning three numbers per image.

AxisMeaningRange
Xhorizontal centroid of the subject−0.5 (left edge) to +0.5 (right edge), 0 = centre
Yvertical centroid−0.5 (top) to +0.5 (bottom), positive is lower in frame
Ssubject size, the square root of the mask's area fraction0 to 1
One recorded limit of the instrument, carried over. On synthetic images its position axes recover the truth almost exactly (worst correlation 0.9979), but for subjects smaller than 8% of the frame the size axis degrades to 0.869. In this corpus the median subject size is 0.48 and only 1.3% of subjects fall below that threshold, so the caveat barely applies here. It is stated because the size axis carries the coupling result in section 7.

The instrument was frozen and hashed in the earlier round, before this engine was ever run, and analyze_flux.py refuses to measure if either file's sha256 differs from that freeze. It matched (appendix A6). Its known weakness is recorded and tested in section 9: where a subject touches the frame edge or fills the frame, the border background model can fit the subject instead of the backdrop.

Protocol

frozen The claim, the kill conditions, the estimand, the referee checks and the decision rules were written before any evaluation image existed, and inherited unchanged from the earlier round: only the engine changed. Changes are dated in an amendments file, each made before measurement. exploratory marks everything computed after the verdict was read; none of it can change a verdict, and it is confined to sections 5 through 8 where it is labelled.

The statistic is a variance share: of the spread in one measurement across prompts and seeds, how much belongs to the prompt. It is an ICC in a crossed prompt-by-seed design, with the McGraw and Wong interval, computed after each subject's own mean is removed so that a subject's habits cannot be counted as prompt control. A cell expected to reach confirms when the lower bound clears 0.5 and fails when the upper bound falls below it. A cell expected not to be reached confirms when the upper bound falls below 0.5 and leaks when the lower bound clears it.

The amendment record

Two amendments, both dated before any evaluation image was measured.

AmendmentWhat changed
Amendment 1, 2026-09-15budget fallback applied (G1 and G3 only, 324 images)
Amendment 2, 2026-09-15the analysis path is made consistent with Amendment 1 (G2 absent)

The second is worth stating plainly because it nearly decided the outcome on a technicality. With the size-named family absent, the inherited analysis crashed outright, its expectation table still listed three cells that could never be measured, and the determinism check re-rendered a cell from the missing family, which would have forced an inconclusive verdict for reasons having nothing to do with the engine. All three were fixed before measurement, and the fix cannot make the verdict more favourable: it only lets the six generated cells be judged at all.

2. How the run was executed

This section is operational rather than scientific, and it is here because the run had already destroyed itself twice before it succeeded.

Two kernel panics during earlier work traced to the same cause: sustained memory pressure with the swap file growing against a nearly full disk, reported by the system as a watchdog timeout. The generator was therefore restructured to run one seed per chunk, stopping and restarting the inference server between chunks so memory is released rather than accumulating across a long run, with a supervisor polling every 60 seconds that halts cleanly if free disk falls below 12 GB or swap passes 10 GB. Finished images are recorded in a manifest and skipped on resume, so any halt is safe.

TimeSeedEvent
21:06:1150000already complete (54/54)
21:06:1115472941/54 done, disk 25.7 GB free, swap 3.2 GB
21:25:27154729chunk ended, 54/54 done, disk 22.5 GB free, swap 5.8 GB
21:25:472594580/54 done, disk 22.5 GB free, swap 5.8 GB
22:42:02259458chunk ended, 54/54 done, disk 21.2 GB free, swap 6.4 GB
22:42:223641870/54 done, disk 21.2 GB free, swap 6.4 GB
23:56:38364187chunk ended, 54/54 done, disk 20.4 GB free, swap 6.3 GB
23:56:584689160/54 done, disk 20.4 GB free, swap 6.3 GB
01:12:13468916chunk ended, 54/54 done, disk 18.0 GB free, swap 7.2 GB
01:12:335736450/54 done, disk 18.0 GB free, swap 7.2 GB
02:26:48573645chunk ended, 54/54 done, disk 17.6 GB free, swap 7.4 GB

The six chunks ran 7.9 hours of inference across a 324-image corpus at a median of 82.8 s per image. Swap rose within each chunk and fell back at every restart, which is the behaviour the design was built to produce. The run survived the controlling session ending partway through, because the generator was detached from it.

Render time per image across the whole run. The regular steps are chunk boundaries, where the server was stopped and the model reloaded.

3. The verdict

KILLED (failure). Two of the six cells failed. Prompts that name a position do not control position on this engine. Every referee check passed, so this is a measured failure of the claim, not an inconclusive result.
FamilyAxisExpectedP95% intervalQPromptsResult
position namedhorizontal positionreach0.071[0.000, 0.254]0.2021FAILS
position namedvertical positionreach0.182[0.048, 0.407]0.1221FAILS
position namedsizenot0.072[0.000, 0.238]0.3421confirms
plain promptshorizontal positionnot0.000[0.000, 0.045]0.4625confirms
plain promptsvertical positionnot0.115[0.015, 0.288]0.2325confirms
plain promptssizenot0.106[0.018, 0.262]0.5425confirms
The six cells. Blue confirms the prediction, red fails it. The dashed line is the 0.5 decision threshold. Both position cells sit entirely below it.

The two failures are unambiguous. Horizontal position under position-named prompts reaches P = 0.071 with an interval of [0.000, 0.254], so the whole interval lies below the threshold; vertical position reaches 0.182 [0.048, 0.407]. The four remaining cells confirm that unnamed axes are not reached, which they do here for an uninteresting reason: on this engine the named axes are not reached either.

What I predicted, before any image existed

The frozen predictions of the bet itself are the six cells above. Separately, and recorded in the same file, I wrote down what I expected this engine to do differently, so that a surprise would be recognisable as one. I was wrong, and in the most useful direction.

CellWhat I expected for FLUXWhat happenedRight?
G1 horizontalconfirms, similar or higherFAILS, P 0.071no
G1 verticalhigher than Z-Image, still possibly openFAILS, P 0.182no
G1 sizeopen or leakconfirms, P 0.072yes
G2 sizeopennot generated (Amendment 1)—
G2 / G3 positionconfirms not reachedX confirms, Y confirmsyes
G3 sizeopenconfirms, P 0.106yes

The stated overall expectation was: “the verdict is INCONCLUSIVE again, with the horizontal-position control replicating. The most informative outcome would be vertical position confirming on FLUX, which would make Z-Image's resistance to "high" an engine property rather than a general one.” Neither half held. The verdict is not inconclusive, and the horizontal control did not replicate, it failed. The reasoning behind that prediction, that a T5 text encoder follows spatial language literally, is exactly the kind of plausible prior this design exists to overrule.

Referee checks

CheckWhat it testsResult
R1instrument on synthetic images with known answerspass, carried over: 400 synthetic images, valid 99.5%, worst correlation 0.9979 across background, greyscale and polarity subgroups
R2interval coverage under subject centringpass, carried over: 400 simulations per cell, worst coverage 0.887 at the chosen 95% level
R3determinism: three cells re-rendered and comparedpass, max abs diff 0, max abs diff 0, max abs diff 0
R4seed crossing: every prompt on the same six seedspass
R5mask validity, limit 5% per familypass, position named 4.32% (7/162), plain 1.23% (2/162)
R6visual audit of 30 mask overlayspass at the declared limit, 6 of 30 wrong (limit 6)
R6 passed at exactly its declared limit, 6 wrong masks of 30 with a limit of 6, and all 6 fall in the position-named family, which is the family that failed. That is the worst possible place for an instrument weakness to concentrate, so it is tested directly in section 9 rather than accepted.

4. The operating window

The verdict says the prompt does not control position. A separate and coarser question is whether the subject ever lands where it was told, which needs no variance decomposition at all: compare the third of the frame the subject occupies with the third that was named. Chance is about 11% on both axes at once.

Instructed positionHit rateMean measured XMean measured YImages
top-left11%-0.040-0.02718
top-centre18%-0.024+0.03717
top-right6%+0.099+0.06816
mid-left6%-0.065+0.10316
centre56%+0.013+0.12818
mid-right6%+0.099+0.20618
bot-left6%+0.003+0.14017
bot-centre72%+0.013+0.23818
bot-right29%+0.095+0.25617
Hit rate by instructed position. Only the centre column functions; the corners and side positions sit at or near chance.

Two instructions work: centre (56%) and bottom-centre (72%). The other seven are at or near chance. The mean measured positions explain why: told top-left, the subject lands at X = -0.040, essentially the centre of the frame; told top-right, +0.099. The engine is not moving the subject far enough for the instruction to register.

Vertically the pattern needs stating carefully, because the obvious reading is wrong. Every instruction naming the top returns a mean Y at or below the centre line (-0.027, +0.037, +0.068 for the three top positions, where negative would mean higher in the frame), which looks like an engine that refuses to raise a subject. It is not. Averaged across the three rows, being told “top” moves the subject up by 0.123 from where the middle instruction puts it, while being told “bottom” moves it down by 0.064: the upward response is the larger of the two. What defeats it is the starting point. The default sits low enough (+0.147) that a shift of 0.123 still lands short of the upper third, so a real response scores as a miss on the hit rate. The bias lives in this engine's default composition, not in its willingness to respond.

Every position-named image by instructed column. Told left or right, the distributions sit almost on top of the centre distribution.

5. Repeatability is not obedience

exploratory This section was written after the verdict.

The statistic and the hit rate measure different things, and the distinction matters for reading any result of this kind. The variance share asks whether a prompt moves the subject to the same place every time. The hit rate asks whether that place is the right one. A model can score high on the first and low on the second by reliably putting the subject somewhere the words did not ask for.

InstructionHit rate on this axisMeasured leftMeasured centreMeasured rightn
told left19.6%1037451
told centre94.3%350053
told right37.3%4281951

Here both are low, which is the cleanest reading of the failure: told left, the subject lands left 19.6% of the time, and sometimes lands on the opposite side entirely (4 images). Overall the subject sits in the instructed third 51.0% of the time horizontally and 50.3% vertically, against 33.3% chance. The engine is responding to something — both figures beat chance — but not enough, and not consistently.

6. Who owns the variance

exploratory

A natural reading of a collapsed prompt share is that the subject stopped moving. That is not what happened. Decomposing the spread of each measurement into what the prompt explains, what the seed explains, and what only the particular pairing of the two explains:

FamilyAxisRaw SDCentred SDPrompt SDSeed SDResidual SDPQResidual share
position namedhorizontal position0.1450.1350.0560.0620.1060.180.210.61
position namedvertical position0.1900.1870.0980.0700.1440.270.140.59
position namedsize0.1580.1610.0640.0930.1140.160.340.51
plain promptshorizontal position0.0540.0530.0140.0360.0370.070.450.49
plain promptsvertical position0.0880.0860.0390.0410.0640.210.230.56
plain promptssize0.1440.1300.0530.0940.0730.160.520.31
These shares are not the same estimator as section 3, and they read higher. The verdict uses ICC(A,1), which corrects for the fact that a prompt's mean is itself estimated from only six seeds and so carries noise of its own; the table above is a plain descriptive split of the observed sums of squares, which credits that noise to the prompt. For horizontal position under position-named prompts the descriptive split gives 0.18 where the corrected estimate is 0.071. The corrected number is the one the verdict rests on. The table is here for the ordering of the three terms, which no correction changes: the residual is the largest.
How the variance in each measurement divides between the prompt, the seed, and the pairing of the two.

For horizontal position under position-named prompts, the prompt and the seed explain 18% and 21%, and the residual explains 61%. The largest single contributor is neither of the two factors: it is the specific combination of one prompt with one seed, which by construction is not predictable from either alone. The subject does move; what it does not do is move anywhere that the words, or the draw, can be said to have chosen.

A share is a ratio, so read it next to the spread it divides. When positions are named the subject's horizontal spread is SD 0.145; under plain prompts it is 0.054, roughly a third as wide. So the seed “owning” 45% of horizontal position in plain prompts is owning a large share of very little movement: with nothing named, this engine puts the subject near the centre and mostly leaves it there. The earlier round taught this lesson on a lighting run, and it applies here unchanged.

With nothing named, the ordering changes. In plain prompts the seed owns 45% of horizontal position and 52% of size, against prompt shares of 7% and 16%. When the words say nothing about geometry, the draw is the strongest thing in the room. That is a hypothesis worth a frozen test of its own; it was measured after the fact and is not a result of this study.

7. Displacement requires shrinkage

exploratory This is the most interesting thing in the corpus, and the one that explains the failure rather than describing it.

Hold the prompt fixed and look only across its six seeds, so the subject, the wording and the instruction are all constant. The renders that pushed the subject furthest from the centre are the renders that made it smallest.

FamilyPromptsr(size, |horizontal|)r(size, distance from centre)
position named27-0.646-0.740
plain prompts27-0.325-0.621
Subject size against distance from centre. Every point is one image; the relationship holds inside a single prompt, not merely across prompts.

The coupling is strong under position-named prompts (r = -0.65 horizontally, -0.74 radially) and survives in plain prompts (r = -0.32, -0.62) where no prompt mentions position or size at all. It is therefore a property of how this generator composes a frame, not a response to instructions.

That reframes the failure. Moving a subject off centre costs frame area: a large subject cannot sit in a corner without leaving the frame. An engine that will not shrink a subject cannot displace it, whatever the words say. At corner prompts specifically, this engine holds subject size at 0.418 and reaches a mean distance of only 0.133 from centre, against 0.475 size at the centre instruction.

The bet contained a hidden assumption. It treated horizontal position, vertical position and size as three axes that a prompt could name independently. On this evidence they are not independent: position and size are coupled through how much of the frame the subject occupies. A bet written on separable axes cannot cleanly describe an engine whose axes are joined, and that is true whether the engine obeys or not.

8. The default composition

exploratory

With nothing named, where does the engine put things?

Every plain-prompt image. The red cross is the frame centre; grey lines mark the thirds.

Subjects cluster near the horizontal centre and sit slightly low: mean X +0.021, mean Y +0.128, with 99% of subjects in the central column. The downward bias is the same direction the instructed positions showed: this engine's idea of a well-composed frame places the subject at or below the middle, never above it.

Subject size under the two prompt families.

Size medians are 0.425 for position-named prompts and 0.522 for plain ones. Naming a position does shrink the subject somewhat, which is the coupling of section 7 acting weakly. It is simply not enough shrinkage to deliver the position.

Subject dependence, and its absence

SubjectHorizontal hit rateMean sizeImages
blue ceramic vase51%0.37851
black cat51%0.47051
man in a red jacket51%0.45853

The three subjects behave almost identically. That flatness is itself informative: when a model is responding to a placement instruction, what is being placed tends to matter, since a rigid object, a deformable animal and a person afford different framings. Uniformity across all three is the signature of an instruction that is not being acted on.

What was dropped, and what that costs

9 of 324 images produced no usable mask, which is 2.8% and comfortably inside the 5% per family the protocol allows. That number understates the effect on the analysis, and the difference is worth being explicit about.

The rule is that a prompt with any invalid image is dropped whole, so that no cell is computed from an uneven set of seeds. Those 9 images therefore removed 6 of 27 position-named prompts and 2 of 27 plain prompts: 36 and 12 images, or 22% of the position-named family rather than the 4.3% the validity rate suggests. That is why the cells above rest on 21 and 25 prompts, not 27.

FamilySubjectPromptInvalid imagesImages removed
position namedblue ceramic vasetop-centre16
position namedblue ceramic vasetop-right16
position namedblue ceramic vasebot-right16
position namedblack catmid-left26
position namedblack catbot-left16
position namedman in a red jackettop-right16
plain promptsblue ceramic vaseparaphrase 1 of 916
plain promptsblack catparaphrase 4 of 916
The dropped prompts are not evenly spread, and this is a threat to validity. Of the 6 position-named prompts lost, five name an off-centre position and none is the centre instruction. The whole-prompt rule therefore removes disproportionately the instructions asking for the largest displacement, which are both the ones the instrument struggles with and the ones a competent engine would have to work hardest to satisfy. The direction of the resulting bias is not obvious: losing hard cases could flatter the engine by removing its worst attempts, or penalise it by removing the cases where it did displace the subject far enough to break the mask. It cannot be settled from this data, and nothing here should be read as though the family were complete.

9. Is it the engine, or the instrument?

This is the question the whole result rests on. The instrument's known weakness is that a subject touching the frame edge or filling the frame can contaminate the border background model, and all 6 wrong masks in the audit fall in the family that failed.

Audit cellWhat is wrong
#4G1 p6 s154729: vase cropped at left edge left unfilled; mask is a crescent of peach background beside it
#6G1 p8 s50000: empty mask, valid=0
#11G1 p18 s154729: red sleeve traced, then mask runs off across white background at top
#12G1 p3 s154729: vase unfilled; large mask blob on the background to its right
#17G1 p20 s364187: man filled but mask also swallows a wide band of background to his left
#27G1 p5 s154729: same failure as #4, vase unfilled, mask on background beside it
The clearest instrument failures, mask filled in magenta. The vase is left unfilled while the mask sits on the background beside it: the border model fitted the backdrop instead of the subject. These are counted as wrong in the audit.
For contrast, three cells counted as NOT wrong. In the first two the extra magenta is the cast shadow on the ground plane; in the third the mask is on the subject but misses part of the beard. The declared rule counts neither as wrong, and the same reading was applied in the earlier round.
All 30 audited masks, filled in magenta, with the six counted as wrong marked in red. The full sheet is reproduced so the count can be checked rather than taken on trust.

Three bearings say the failure is the engine anyway.

First, the images themselves, with no instrument involved. One subject, one seed, all nine instructed positions:

The vase at all nine instructed positions, one seed. Seven of the nine are a large, roughly centred vase regardless of the words.
The same for the man. Left-hand instructions displace him; right-hand ones mostly do not.

Second, the coarse hit rate of section 4 does not use the variance machinery at all, and it agrees: the subject reaches the instructed third 51.0% of the time horizontally.

Third, a structural split fixed by the prompt wording rather than by the results. The instrument struggles with cropped, corner-ish images; if those were driving the failure, prompts that name an edge or the centre rather than a corner should recover.

Prompt typePromptsP95% intervalQ
corner words80.111[0.000, 0.571]0.31
edge and centre words130.107[0.000, 0.416]0.08

They do not recover: 0.111 and 0.107, indistinguishable. The instrument weakness is real, is recorded in the referee file, and does not account for the result.

What the weakness does cost. It inflates the residual term of section 6, since a mismeasured image contributes unrepeatable noise. The precise split between prompt, seed and residual should be read as approximate. The direction of the finding does not depend on it.

10. What this suggests testing next

Three follow-ups, in the order I would run them. The first is a control, the second is the real experiment, and the third is maintenance on the instrument.

  1. Raise the text encoder's precision and change nothing else. The T5 here is quantised to 4 bits, which is a plausible place for spatial language to degrade before the sampler ever sees it. Rerunning the position-named family at higher encoder precision, with identical prompts, seeds and sampler settings, would separate “this model does not follow position words” from “this quantisation does not carry them”. Until that is run, the result belongs to the configuration and not to FLUX.
  2. Test the coupling as its own frozen bet: an engine can displace a subject only in proportion to how much it shrinks it. Name size and position together, and set them against each other: ask for a large subject in a corner, a small subject dead centre, and the compatible pairs. If the coupling is real, the conflicting requests should be resolved in a consistent direction, and that direction is the interesting number. This would explain both rounds instead of describing them, and it is falsifiable in one 324-image run.
  3. Rebuild the instrument for subjects that touch the frame. Both rounds now hit the same wall: a border-ring background model inverts when the subject reaches the edge or fills the frame. A second, independent reader that does not assume the border is background would let the position cells be measured without the caveat in section 9, and would retire the largest standing objection to both studies.

The second is the one worth doing first if only one gets run, because a confirmed coupling would change how the earlier round's results are read as well.

11. Limits, and what this does not prove

12. Files and reproduction

Everything in this paper is rebuilt from the study folder by build_paper.py; no number is typed by hand. The study itself reproduces from its own files.

FileRole
REPLICATION.mdthe frozen plan, written before any image
AMENDMENTS.mdthe two dated amendments, both before measurement
VERDICT.mdthe verdict and its reasoning
INSIGHTS.mdthe exploratory findings, labelled and dated after the verdict
results.json, measurements.csvthe six cells, and 324 measured images
referee_r3.json, referee_r6.jsondeterminism, and the visual audit with its judgement recorded
run_safeguarded.py, run_safeguarded.logthe chunked generator and its record
gen_flux.py, analyze_flux.pygeneration and analysis entry points

Appendix: the recipe

A1. Prompt templates

Each family has 27 prompts: three subjects crossed with nine variants. The position family names one of nine positions; the plain family names nothing geometric. One example per subject per family:

FamilySubjectPrompt
position namedblue ceramic vaseA photograph of a blue ceramic vase in the top-left corner of the frame, with a plain background.
position namedblack catA photograph of a black cat in the top-left corner of the frame, with a plain background.
position namedman in a red jacketA photograph of a man in a red jacket in the top-left corner of the frame, with a plain background.
plain promptsblue ceramic vaseA photograph of a blue ceramic vase, with a plain background.
plain promptsblack catA photograph of a black cat, with a plain background.
plain promptsman in a red jacketA photograph of a man in a red jacket, with a plain background.

The nine position phrases, in order: top-left, top-centre, top-right, mid-left, centre, mid-right, bot-left, bot-centre, bot-right. The nine plain variants are neutral paraphrases (“A photograph of…”, “A clean product-style photo of…”, and so on) that vary the wording without naming geometry.

A2. Engine and sampling

SettingValue
Modelflux1-schnell-Q5_K_S.gguf, 8,263,222,304 bytes
Text encoderst5-v1_1-xxl-encoder-Q4_K_M.gguf and clip_l.safetensors
VAEae.safetensors
Resolution768 × 768
Steps / cfg4 / 1.0
Sampler / scheduleeuler / simple
Negative conditioningzeroed (cfg 1 makes it inert)
Seeds50000, 154729, 259458, 364187, 468916, 573645
HostApple M1 Max, 32 GB, headless ComfyUI 0.17.2
Median render82.8 s per image

A3. What each axis means

The mask's centroid gives X and Y in frame coordinates, with the origin at the centre and the frame spanning −0.5 to +0.5 on each side. S is the square root of the mask's area fraction, so it scales like a length rather than an area. A cell of the study is one family crossed with one axis, and each cell is judged on its own.

A4. The statistic

For one cell, arrange the measurements as a prompts-by-seeds matrix, subtract each subject's own mean, and take the intraclass correlation ICC(A,1) in the crossed design: the share of variance attributable to the prompt, with the McGraw and Wong F-based 95% interval and degrees of freedom reduced by the number of subject means removed. Q is the matching share for the seed. The residual share reported in section 6 is what remains, the prompt-by-seed interaction together with measurement noise.

A5. Referee checks

CheckRule as declared
R1the instrument recovers known positions and sizes on synthetic images; carried over, engine-independent
R2the interval covers at its nominal rate under subject centring; carried over, engine-independent
R3at least three cells re-rendered from the same seed match within a max absolute pixel difference of 2
R4every prompt in a family is generated on the same set of six seeds
R5no more than 5% of a family's images yield an invalid mask; a prompt with any invalid mask is dropped whole
R6at most 6 of 30 randomly drawn mask overlays visibly wrong (mask on background, missing the subject, or merged with a horizon)

R1 and R2 test the instrument and the statistic rather than the engine, so their results were carried over from the round in which they were established, and recorded as carried over rather than re-run.

A6. Hashes and commands

Filesha256
objmask.py3e4bc7a3dde1cca5ebcf33904e3c56e8…
stats_centred.py99eb21644a164d3e379997d2bf4a9edc…
gen2.pybd03d34c5566d5e71ef894c7a721351c…
../seed-bet/gen.pya401f5802e088b3d0662b5e159b99f59…
objmask.py at measurement3e4bc7a3dde1cca5ebcf33904e3c56e8…
stats_centred.py at measurement99eb21644a164d3e379997d2bf4a9edc…

The analysis refuses to run if the first two differ from the freeze. Reproduction, in order:

python3 run_safeguarded.py --families G1,G3 · python3 gen_flux.py rerender · python3 analyze_flux.py r3 · python3 analyze_flux.py measure · python3 analyze_flux.py audit · python3 analyze_flux.py analyze

A7. Terms

TermMeaning here
Cellone prompt family crossed with one measured axis; six in this study
P, prompt sharefraction of variance in a measurement attributable to which prompt was used
Q, seed sharethe same, for which seed was used
Reacha cell where the bet predicts the prompt controls the axis
Leakan unnamed axis that the prompt turns out to control anyway
Hit ratefraction of images where the subject occupies the instructed third; unrelated to P
Frozen / exploratoryfixed before any evaluation image / computed after the verdict was read