A bet was written down before any image existed: within a subject, a prompt controls exactly the geometric axes it names, and leaves the others to the sampler's draw. On the engine it was written for it came through without a single cell failing. Run unchanged on FLUX.1-schnell, it died. Position words do not place the subject here. This paper reports what happened instead, and what the failure exposes about the assumption underneath the bet: that position and size are separate axes at all.
- Setup: engine, corpus, instrument, protocol
- How the run was executed
- The verdict
- The operating window
- Repeatability is not obedience
- Who owns the variance
- Displacement requires shrinkage
- The default composition
- Is it the engine, or the instrument?
- What this suggests testing next
- Limits, and what this does not prove
- Files and reproduction
- Appendix: the recipe
1. Setup
The engine and the corpus
Every image comes from FLUX.1-schnell, quantised to Q5_K_S GGUF (8.26 GB, city96 build,
Apache-2.0), with a 4-bit T5-XXL text encoder and CLIP-L, running locally in headless ComfyUI 0.17.2 on an
Apple M1 Max with 32 GB. Images are 768 × 768, 4 steps, cfg 1, euler with the
simple schedule. Generation is deterministic: three re-rendered cells came back pixel-identical
to their originals (R3, section 3).
| Set | Prompts | Seeds | Images | Purpose |
|---|---|---|---|---|
| Position named | 27 (3 subjects × 9 instructed positions) | 6 | 162 | does a position word place the subject? |
| Plain prompts | 27 (9 neutral paraphrases × 3 subjects) | 6 | 162 | the engine's default composition |
The same six seeds run across every prompt, so a seed's own effect is separable from the prompt's. The three subjects are a blue ceramic vase, a black cat, and a man in a red jacket: one rigid object, one deformable animal, one person.
The instrument
One reader measures every image: an object-level mask that finds the subject against a background model fitted to the frame border, returning three numbers per image.
| Axis | Meaning | Range |
|---|---|---|
| X | horizontal centroid of the subject | −0.5 (left edge) to +0.5 (right edge), 0 = centre |
| Y | vertical centroid | −0.5 (top) to +0.5 (bottom), positive is lower in frame |
| S | subject size, the square root of the mask's area fraction | 0 to 1 |
The instrument was frozen and hashed in the earlier round, before this engine was ever run, and
analyze_flux.py refuses to measure if either file's sha256 differs from that freeze. It matched
(appendix A6). Its known weakness is recorded and tested in section 9: where a subject touches the frame edge
or fills the frame, the border background model can fit the subject instead of the backdrop.
Protocol
frozen The claim, the kill conditions, the estimand, the referee checks and the decision rules were written before any evaluation image existed, and inherited unchanged from the earlier round: only the engine changed. Changes are dated in an amendments file, each made before measurement. exploratory marks everything computed after the verdict was read; none of it can change a verdict, and it is confined to sections 5 through 8 where it is labelled.
The statistic is a variance share: of the spread in one measurement across prompts and seeds, how much belongs to the prompt. It is an ICC in a crossed prompt-by-seed design, with the McGraw and Wong interval, computed after each subject's own mean is removed so that a subject's habits cannot be counted as prompt control. A cell expected to reach confirms when the lower bound clears 0.5 and fails when the upper bound falls below it. A cell expected not to be reached confirms when the upper bound falls below 0.5 and leaks when the lower bound clears it.
The amendment record
Two amendments, both dated before any evaluation image was measured.
| Amendment | What changed |
|---|---|
| Amendment 1, 2026-09-15 | budget fallback applied (G1 and G3 only, 324 images) |
| Amendment 2, 2026-09-15 | the analysis path is made consistent with Amendment 1 (G2 absent) |
The second is worth stating plainly because it nearly decided the outcome on a technicality. With the size-named family absent, the inherited analysis crashed outright, its expectation table still listed three cells that could never be measured, and the determinism check re-rendered a cell from the missing family, which would have forced an inconclusive verdict for reasons having nothing to do with the engine. All three were fixed before measurement, and the fix cannot make the verdict more favourable: it only lets the six generated cells be judged at all.
2. How the run was executed
This section is operational rather than scientific, and it is here because the run had already destroyed itself twice before it succeeded.
Two kernel panics during earlier work traced to the same cause: sustained memory pressure with the swap file growing against a nearly full disk, reported by the system as a watchdog timeout. The generator was therefore restructured to run one seed per chunk, stopping and restarting the inference server between chunks so memory is released rather than accumulating across a long run, with a supervisor polling every 60 seconds that halts cleanly if free disk falls below 12 GB or swap passes 10 GB. Finished images are recorded in a manifest and skipped on resume, so any halt is safe.
| Time | Seed | Event |
|---|---|---|
| 21:06:11 | 50000 | already complete (54/54) |
| 21:06:11 | 154729 | 41/54 done, disk 25.7 GB free, swap 3.2 GB |
| 21:25:27 | 154729 | chunk ended, 54/54 done, disk 22.5 GB free, swap 5.8 GB |
| 21:25:47 | 259458 | 0/54 done, disk 22.5 GB free, swap 5.8 GB |
| 22:42:02 | 259458 | chunk ended, 54/54 done, disk 21.2 GB free, swap 6.4 GB |
| 22:42:22 | 364187 | 0/54 done, disk 21.2 GB free, swap 6.4 GB |
| 23:56:38 | 364187 | chunk ended, 54/54 done, disk 20.4 GB free, swap 6.3 GB |
| 23:56:58 | 468916 | 0/54 done, disk 20.4 GB free, swap 6.3 GB |
| 01:12:13 | 468916 | chunk ended, 54/54 done, disk 18.0 GB free, swap 7.2 GB |
| 01:12:33 | 573645 | 0/54 done, disk 18.0 GB free, swap 7.2 GB |
| 02:26:48 | 573645 | chunk ended, 54/54 done, disk 17.6 GB free, swap 7.4 GB |
The six chunks ran 7.9 hours of inference across a 324-image corpus at a median of 82.8 s per image. Swap rose within each chunk and fell back at every restart, which is the behaviour the design was built to produce. The run survived the controlling session ending partway through, because the generator was detached from it.
3. The verdict
| Family | Axis | Expected | P | 95% interval | Q | Prompts | Result |
|---|---|---|---|---|---|---|---|
| position named | horizontal position | reach | 0.071 | [0.000, 0.254] | 0.20 | 21 | FAILS |
| position named | vertical position | reach | 0.182 | [0.048, 0.407] | 0.12 | 21 | FAILS |
| position named | size | not | 0.072 | [0.000, 0.238] | 0.34 | 21 | confirms |
| plain prompts | horizontal position | not | 0.000 | [0.000, 0.045] | 0.46 | 25 | confirms |
| plain prompts | vertical position | not | 0.115 | [0.015, 0.288] | 0.23 | 25 | confirms |
| plain prompts | size | not | 0.106 | [0.018, 0.262] | 0.54 | 25 | confirms |
The two failures are unambiguous. Horizontal position under position-named prompts reaches P = 0.071 with an interval of [0.000, 0.254], so the whole interval lies below the threshold; vertical position reaches 0.182 [0.048, 0.407]. The four remaining cells confirm that unnamed axes are not reached, which they do here for an uninteresting reason: on this engine the named axes are not reached either.
What I predicted, before any image existed
The frozen predictions of the bet itself are the six cells above. Separately, and recorded in the same file, I wrote down what I expected this engine to do differently, so that a surprise would be recognisable as one. I was wrong, and in the most useful direction.
| Cell | What I expected for FLUX | What happened | Right? |
|---|---|---|---|
| G1 horizontal | confirms, similar or higher | FAILS, P 0.071 | no |
| G1 vertical | higher than Z-Image, still possibly open | FAILS, P 0.182 | no |
| G1 size | open or leak | confirms, P 0.072 | yes |
| G2 size | open | not generated (Amendment 1) | — |
| G2 / G3 position | confirms not reached | X confirms, Y confirms | yes |
| G3 size | open | confirms, P 0.106 | yes |
The stated overall expectation was: “the verdict is INCONCLUSIVE again, with the horizontal-position control replicating. The most informative outcome would be vertical position confirming on FLUX, which would make Z-Image's resistance to "high" an engine property rather than a general one.” Neither half held. The verdict is not inconclusive, and the horizontal control did not replicate, it failed. The reasoning behind that prediction, that a T5 text encoder follows spatial language literally, is exactly the kind of plausible prior this design exists to overrule.
Referee checks
| Check | What it tests | Result |
|---|---|---|
| R1 | instrument on synthetic images with known answers | pass, carried over: 400 synthetic images, valid 99.5%, worst correlation 0.9979 across background, greyscale and polarity subgroups |
| R2 | interval coverage under subject centring | pass, carried over: 400 simulations per cell, worst coverage 0.887 at the chosen 95% level |
| R3 | determinism: three cells re-rendered and compared | pass, max abs diff 0, max abs diff 0, max abs diff 0 |
| R4 | seed crossing: every prompt on the same six seeds | pass |
| R5 | mask validity, limit 5% per family | pass, position named 4.32% (7/162), plain 1.23% (2/162) |
| R6 | visual audit of 30 mask overlays | pass at the declared limit, 6 of 30 wrong (limit 6) |
4. The operating window
The verdict says the prompt does not control position. A separate and coarser question is whether the subject ever lands where it was told, which needs no variance decomposition at all: compare the third of the frame the subject occupies with the third that was named. Chance is about 11% on both axes at once.
| Instructed position | Hit rate | Mean measured X | Mean measured Y | Images |
|---|---|---|---|---|
| top-left | 11% | -0.040 | -0.027 | 18 |
| top-centre | 18% | -0.024 | +0.037 | 17 |
| top-right | 6% | +0.099 | +0.068 | 16 |
| mid-left | 6% | -0.065 | +0.103 | 16 |
| centre | 56% | +0.013 | +0.128 | 18 |
| mid-right | 6% | +0.099 | +0.206 | 18 |
| bot-left | 6% | +0.003 | +0.140 | 17 |
| bot-centre | 72% | +0.013 | +0.238 | 18 |
| bot-right | 29% | +0.095 | +0.256 | 17 |
Two instructions work: centre (56%) and bottom-centre (72%). The other seven are at or near chance. The mean measured positions explain why: told top-left, the subject lands at X = -0.040, essentially the centre of the frame; told top-right, +0.099. The engine is not moving the subject far enough for the instruction to register.
Vertically the pattern needs stating carefully, because the obvious reading is wrong. Every instruction naming the top returns a mean Y at or below the centre line (-0.027, +0.037, +0.068 for the three top positions, where negative would mean higher in the frame), which looks like an engine that refuses to raise a subject. It is not. Averaged across the three rows, being told “top” moves the subject up by 0.123 from where the middle instruction puts it, while being told “bottom” moves it down by 0.064: the upward response is the larger of the two. What defeats it is the starting point. The default sits low enough (+0.147) that a shift of 0.123 still lands short of the upper third, so a real response scores as a miss on the hit rate. The bias lives in this engine's default composition, not in its willingness to respond.
5. Repeatability is not obedience
exploratory This section was written after the verdict.
The statistic and the hit rate measure different things, and the distinction matters for reading any result of this kind. The variance share asks whether a prompt moves the subject to the same place every time. The hit rate asks whether that place is the right one. A model can score high on the first and low on the second by reliably putting the subject somewhere the words did not ask for.
| Instruction | Hit rate on this axis | Measured left | Measured centre | Measured right | n |
|---|---|---|---|---|---|
| told left | 19.6% | 10 | 37 | 4 | 51 |
| told centre | 94.3% | 3 | 50 | 0 | 53 |
| told right | 37.3% | 4 | 28 | 19 | 51 |
Here both are low, which is the cleanest reading of the failure: told left, the subject lands left 19.6% of the time, and sometimes lands on the opposite side entirely (4 images). Overall the subject sits in the instructed third 51.0% of the time horizontally and 50.3% vertically, against 33.3% chance. The engine is responding to something — both figures beat chance — but not enough, and not consistently.
6. Who owns the variance
exploratory
A natural reading of a collapsed prompt share is that the subject stopped moving. That is not what happened. Decomposing the spread of each measurement into what the prompt explains, what the seed explains, and what only the particular pairing of the two explains:
| Family | Axis | Raw SD | Centred SD | Prompt SD | Seed SD | Residual SD | P | Q | Residual share |
|---|---|---|---|---|---|---|---|---|---|
| position named | horizontal position | 0.145 | 0.135 | 0.056 | 0.062 | 0.106 | 0.18 | 0.21 | 0.61 |
| position named | vertical position | 0.190 | 0.187 | 0.098 | 0.070 | 0.144 | 0.27 | 0.14 | 0.59 |
| position named | size | 0.158 | 0.161 | 0.064 | 0.093 | 0.114 | 0.16 | 0.34 | 0.51 |
| plain prompts | horizontal position | 0.054 | 0.053 | 0.014 | 0.036 | 0.037 | 0.07 | 0.45 | 0.49 |
| plain prompts | vertical position | 0.088 | 0.086 | 0.039 | 0.041 | 0.064 | 0.21 | 0.23 | 0.56 |
| plain prompts | size | 0.144 | 0.130 | 0.053 | 0.094 | 0.073 | 0.16 | 0.52 | 0.31 |
For horizontal position under position-named prompts, the prompt and the seed explain 18% and 21%, and the residual explains 61%. The largest single contributor is neither of the two factors: it is the specific combination of one prompt with one seed, which by construction is not predictable from either alone. The subject does move; what it does not do is move anywhere that the words, or the draw, can be said to have chosen.
With nothing named, the ordering changes. In plain prompts the seed owns 45% of horizontal position and 52% of size, against prompt shares of 7% and 16%. When the words say nothing about geometry, the draw is the strongest thing in the room. That is a hypothesis worth a frozen test of its own; it was measured after the fact and is not a result of this study.
7. Displacement requires shrinkage
exploratory This is the most interesting thing in the corpus, and the one that explains the failure rather than describing it.
Hold the prompt fixed and look only across its six seeds, so the subject, the wording and the instruction are all constant. The renders that pushed the subject furthest from the centre are the renders that made it smallest.
| Family | Prompts | r(size, |horizontal|) | r(size, distance from centre) |
|---|---|---|---|
| position named | 27 | -0.646 | -0.740 |
| plain prompts | 27 | -0.325 | -0.621 |
The coupling is strong under position-named prompts (r = -0.65 horizontally, -0.74 radially) and survives in plain prompts (r = -0.32, -0.62) where no prompt mentions position or size at all. It is therefore a property of how this generator composes a frame, not a response to instructions.
That reframes the failure. Moving a subject off centre costs frame area: a large subject cannot sit in a corner without leaving the frame. An engine that will not shrink a subject cannot displace it, whatever the words say. At corner prompts specifically, this engine holds subject size at 0.418 and reaches a mean distance of only 0.133 from centre, against 0.475 size at the centre instruction.
8. The default composition
exploratory
With nothing named, where does the engine put things?
Subjects cluster near the horizontal centre and sit slightly low: mean X +0.021, mean Y +0.128, with 99% of subjects in the central column. The downward bias is the same direction the instructed positions showed: this engine's idea of a well-composed frame places the subject at or below the middle, never above it.
Size medians are 0.425 for position-named prompts and 0.522 for plain ones. Naming a position does shrink the subject somewhat, which is the coupling of section 7 acting weakly. It is simply not enough shrinkage to deliver the position.
Subject dependence, and its absence
| Subject | Horizontal hit rate | Mean size | Images |
|---|---|---|---|
| blue ceramic vase | 51% | 0.378 | 51 |
| black cat | 51% | 0.470 | 51 |
| man in a red jacket | 51% | 0.458 | 53 |
The three subjects behave almost identically. That flatness is itself informative: when a model is responding to a placement instruction, what is being placed tends to matter, since a rigid object, a deformable animal and a person afford different framings. Uniformity across all three is the signature of an instruction that is not being acted on.
What was dropped, and what that costs
9 of 324 images produced no usable mask, which is 2.8% and comfortably inside the 5% per family the protocol allows. That number understates the effect on the analysis, and the difference is worth being explicit about.
The rule is that a prompt with any invalid image is dropped whole, so that no cell is computed from an uneven set of seeds. Those 9 images therefore removed 6 of 27 position-named prompts and 2 of 27 plain prompts: 36 and 12 images, or 22% of the position-named family rather than the 4.3% the validity rate suggests. That is why the cells above rest on 21 and 25 prompts, not 27.
| Family | Subject | Prompt | Invalid images | Images removed |
|---|---|---|---|---|
| position named | blue ceramic vase | top-centre | 1 | 6 |
| position named | blue ceramic vase | top-right | 1 | 6 |
| position named | blue ceramic vase | bot-right | 1 | 6 |
| position named | black cat | mid-left | 2 | 6 |
| position named | black cat | bot-left | 1 | 6 |
| position named | man in a red jacket | top-right | 1 | 6 |
| plain prompts | blue ceramic vase | paraphrase 1 of 9 | 1 | 6 |
| plain prompts | black cat | paraphrase 4 of 9 | 1 | 6 |
9. Is it the engine, or the instrument?
This is the question the whole result rests on. The instrument's known weakness is that a subject touching the frame edge or filling the frame can contaminate the border background model, and all 6 wrong masks in the audit fall in the family that failed.
| Audit cell | What is wrong |
|---|---|
| #4 | G1 p6 s154729: vase cropped at left edge left unfilled; mask is a crescent of peach background beside it |
| #6 | G1 p8 s50000: empty mask, valid=0 |
| #11 | G1 p18 s154729: red sleeve traced, then mask runs off across white background at top |
| #12 | G1 p3 s154729: vase unfilled; large mask blob on the background to its right |
| #17 | G1 p20 s364187: man filled but mask also swallows a wide band of background to his left |
| #27 | G1 p5 s154729: same failure as #4, vase unfilled, mask on background beside it |
Three bearings say the failure is the engine anyway.
First, the images themselves, with no instrument involved. One subject, one seed, all nine instructed positions:
Second, the coarse hit rate of section 4 does not use the variance machinery at all, and it agrees: the subject reaches the instructed third 51.0% of the time horizontally.
Third, a structural split fixed by the prompt wording rather than by the results. The instrument struggles with cropped, corner-ish images; if those were driving the failure, prompts that name an edge or the centre rather than a corner should recover.
| Prompt type | Prompts | P | 95% interval | Q |
|---|---|---|---|---|
| corner words | 8 | 0.111 | [0.000, 0.571] | 0.31 |
| edge and centre words | 13 | 0.107 | [0.000, 0.416] | 0.08 |
They do not recover: 0.111 and 0.107, indistinguishable. The instrument weakness is real, is recorded in the referee file, and does not account for the result.
10. What this suggests testing next
Three follow-ups, in the order I would run them. The first is a control, the second is the real experiment, and the third is maintenance on the instrument.
- Raise the text encoder's precision and change nothing else. The T5 here is quantised to 4 bits, which is a plausible place for spatial language to degrade before the sampler ever sees it. Rerunning the position-named family at higher encoder precision, with identical prompts, seeds and sampler settings, would separate “this model does not follow position words” from “this quantisation does not carry them”. Until that is run, the result belongs to the configuration and not to FLUX.
- Test the coupling as its own frozen bet: an engine can displace a subject only in proportion to how much it shrinks it. Name size and position together, and set them against each other: ask for a large subject in a corner, a small subject dead centre, and the compatible pairs. If the coupling is real, the conflicting requests should be resolved in a consistent direction, and that direction is the interesting number. This would explain both rounds instead of describing them, and it is falsifiable in one 324-image run.
- Rebuild the instrument for subjects that touch the frame. Both rounds now hit the same wall: a border-ring background model inverts when the subject reaches the edge or fills the frame. A second, independent reader that does not assume the border is background would let the position cells be measured without the caveat in section 9, and would retire the largest standing objection to both studies.
The second is the one worth doing first if only one gets run, because a confirmed coupling would change how the earlier round's results are read as well.
11. Limits, and what this does not prove
- One operating point, not a model family. This is
schnell, quantised to 5 bits, at 4 steps and cfg 1, 768 × 768, with a 4-bit text encoder. Any of those could matter. The result says what this configuration does, and a configuration is what people actually run. - The text encoder is quantised too. A 4-bit T5 is a plausible place for spatial language to degrade, and this study cannot separate encoder degradation from sampler behaviour. That is the single most valuable follow-up: rerun at higher encoder precision, change nothing else.
- Six cells, not nine. The size-named family was dropped by a budget rule fixed in advance, so nothing here tests whether size words set size on this engine.
- Three subjects, one framing vocabulary. All prompts are studio-style single-subject descriptions on plain backgrounds. Scene prompts were not run on this engine.
- One instrument. Every number traces to a single object mask with a known and demonstrated failure mode. The bearings in section 9 reduce but do not eliminate that dependence.
- Exploratory sections are not results. Sections 5 through 8 were computed after the verdict was read and cannot support a claim on their own. The coupling in section 7 is stated as the next thing to test, not as a finding.
12. Files and reproduction
Everything in this paper is rebuilt from the study folder by build_paper.py; no number is
typed by hand. The study itself reproduces from its own files.
| File | Role |
|---|---|
REPLICATION.md | the frozen plan, written before any image |
AMENDMENTS.md | the two dated amendments, both before measurement |
VERDICT.md | the verdict and its reasoning |
INSIGHTS.md | the exploratory findings, labelled and dated after the verdict |
results.json, measurements.csv | the six cells, and 324 measured images |
referee_r3.json, referee_r6.json | determinism, and the visual audit with its judgement recorded |
run_safeguarded.py, run_safeguarded.log | the chunked generator and its record |
gen_flux.py, analyze_flux.py | generation and analysis entry points |
Appendix: the recipe
A1. Prompt templates
Each family has 27 prompts: three subjects crossed with nine variants. The position family names one of nine positions; the plain family names nothing geometric. One example per subject per family:
| Family | Subject | Prompt |
|---|---|---|
| position named | blue ceramic vase | A photograph of a blue ceramic vase in the top-left corner of the frame, with a plain background. |
| position named | black cat | A photograph of a black cat in the top-left corner of the frame, with a plain background. |
| position named | man in a red jacket | A photograph of a man in a red jacket in the top-left corner of the frame, with a plain background. |
| plain prompts | blue ceramic vase | A photograph of a blue ceramic vase, with a plain background. |
| plain prompts | black cat | A photograph of a black cat, with a plain background. |
| plain prompts | man in a red jacket | A photograph of a man in a red jacket, with a plain background. |
The nine position phrases, in order: top-left, top-centre, top-right, mid-left, centre, mid-right, bot-left, bot-centre, bot-right. The nine plain variants are neutral paraphrases (“A photograph of…”, “A clean product-style photo of…”, and so on) that vary the wording without naming geometry.
A2. Engine and sampling
| Setting | Value |
|---|---|
| Model | flux1-schnell-Q5_K_S.gguf, 8,263,222,304 bytes |
| Text encoders | t5-v1_1-xxl-encoder-Q4_K_M.gguf and clip_l.safetensors |
| VAE | ae.safetensors |
| Resolution | 768 × 768 |
| Steps / cfg | 4 / 1.0 |
| Sampler / schedule | euler / simple |
| Negative conditioning | zeroed (cfg 1 makes it inert) |
| Seeds | 50000, 154729, 259458, 364187, 468916, 573645 |
| Host | Apple M1 Max, 32 GB, headless ComfyUI 0.17.2 |
| Median render | 82.8 s per image |
A3. What each axis means
The mask's centroid gives X and Y in frame coordinates, with the origin at the centre and the frame spanning −0.5 to +0.5 on each side. S is the square root of the mask's area fraction, so it scales like a length rather than an area. A cell of the study is one family crossed with one axis, and each cell is judged on its own.
A4. The statistic
For one cell, arrange the measurements as a prompts-by-seeds matrix, subtract each subject's own mean, and take the intraclass correlation ICC(A,1) in the crossed design: the share of variance attributable to the prompt, with the McGraw and Wong F-based 95% interval and degrees of freedom reduced by the number of subject means removed. Q is the matching share for the seed. The residual share reported in section 6 is what remains, the prompt-by-seed interaction together with measurement noise.
A5. Referee checks
| Check | Rule as declared |
|---|---|
| R1 | the instrument recovers known positions and sizes on synthetic images; carried over, engine-independent |
| R2 | the interval covers at its nominal rate under subject centring; carried over, engine-independent |
| R3 | at least three cells re-rendered from the same seed match within a max absolute pixel difference of 2 |
| R4 | every prompt in a family is generated on the same set of six seeds |
| R5 | no more than 5% of a family's images yield an invalid mask; a prompt with any invalid mask is dropped whole |
| R6 | at most 6 of 30 randomly drawn mask overlays visibly wrong (mask on background, missing the subject, or merged with a horizon) |
R1 and R2 test the instrument and the statistic rather than the engine, so their results were carried over from the round in which they were established, and recorded as carried over rather than re-run.
A6. Hashes and commands
| File | sha256 |
|---|---|
objmask.py | 3e4bc7a3dde1cca5ebcf33904e3c56e8… |
stats_centred.py | 99eb21644a164d3e379997d2bf4a9edc… |
gen2.py | bd03d34c5566d5e71ef894c7a721351c… |
../seed-bet/gen.py | a401f5802e088b3d0662b5e159b99f59… |
| objmask.py at measurement | 3e4bc7a3dde1cca5ebcf33904e3c56e8… |
| stats_centred.py at measurement | 99eb21644a164d3e379997d2bf4a9edc… |
The analysis refuses to run if the first two differ from the freeze. Reproduction, in order:
python3 run_safeguarded.py --families G1,G3 ·
python3 gen_flux.py rerender ·
python3 analyze_flux.py r3 ·
python3 analyze_flux.py measure ·
python3 analyze_flux.py audit ·
python3 analyze_flux.py analyze
A7. Terms
| Term | Meaning here |
|---|---|
| Cell | one prompt family crossed with one measured axis; six in this study |
| P, prompt share | fraction of variance in a measurement attributable to which prompt was used |
| Q, seed share | the same, for which seed was used |
| Reach | a cell where the bet predicts the prompt controls the axis |
| Leak | an unnamed axis that the prompt turns out to control anyway |
| Hit rate | fraction of images where the subject occupies the instructed third; unrelated to P |
| Frozen / exploratory | fixed before any evaluation image / computed after the verdict was read |