# Stage 1 / v0.10 — completed information-value diagnostic

**Readings sometimes improve care, but this study is inconclusive about whether the present investigation menu offers enough benefit to justify a new selector.** Ensemble care remains the default. No revised selector was trained or promoted.

The [frozen method](v10-method.md) separated the effect of receiving a test's readings from the result of taking the whole test. Twelve fresh habitat seeds supplied five conditions and three clock-sampled checkpoints each. All 180 checkpoints completed. There were 719 eligible paired tests; one isolation prefix exceeded the existing 24-action limit and was excluded before comparison.

## The decision, using the thresholds set before the run

A useful test had to improve both total care over ordinary control and care relative to withholding its readings by at least one percentage point over a 96-tick window.

| Endpoint | Observed | 95% interval across 12 seeds | Declared target |
|---|---:|---:|---:|
| Checkpoints with a useful test | 28/180 / 15.56% | 9.48–21.64% | 20% |
| Mean retrospectively attainable gain from qualifying tests | 0.85 percentage points | 0.16–1.54 pp | 1.00 pp |

Both intervals cross their targets. The frozen rule therefore returns **inconclusive**, rather than sufficient or limited headroom. These are seed-cluster intervals: conditions and moments within a seed are correlated. They are descriptive decision aids, not multiplicity-adjusted hypothesis tests.

The attainable gain is chosen retrospectively among the tested alternatives, with zero assigned when none qualify. It is not a gain achieved by a policy, and it is not an optimal upper bound. The development gate had failed at 2/30 checkpoints; the fresh sample does not retroactively turn development into validation or justify changing its threshold.

## What the decomposition revealed

| Among 719 paired tests | Count |
|---|---:|
| Reading benefit and total benefit both met the threshold | 31 |
| Readings helped relative to withholding, but the complete test did not beat ordinary care by the threshold | 69 |
| Complete test helped, but its reading contribution was below the threshold | 48 |
| Neither benefit met the threshold | 571 |

Receiving the readings changed the first following care choice in 226/719 tests. A changed choice alone is not evidence of improved care. Benefits can also appear later: the saved pipe example begins with the same next journey in both branches, then takes different actions one tick later.

Fourteen of the 28 useful checkpoints occurred during suspected mismatches, and fourteen occurred outside that trigger. There were 51 suspected-mismatch checkpoints overall. This suggests investigating when a reading could change an imminent decision, but it does not show that simply removing the trigger would improve a selector.

The full-state reference improved care by at least one point at 34/180 checkpoints. It uses the same bounded planning machinery with privileged current state and coefficients. It is neither an optimal controller nor proof that additional sensing would deliver those gains.

| Condition | Checkpoints offering a useful test |
|---|---:|
| Steady | 4/36 |
| Pipe fault | 4/36 |
| Drying | 7/36 |
| Supply shortage | 9/36 |
| Compound fault | 4/36 |

The remaining effect in the decomposition includes the physical intervention, route, delayed care and the withholding condition. It must not be read as an isolated physical causal effect. Reading value is conditional on these channels, the withholding window and this controller.

## What is usable now

**Test the reading**, beside the nursery controls, pauses the live habitat and compares ordinary care with the same test with and without its readings. It displays:

- Three synchronized continuations, a shared time slider and play/pause.
- Separate total gain, reading contribution and remaining effect.
- The first following care choice and the first action where the branches diverge.
- A **Jump to first difference** control, observed costs and a clear observer-only boundary.
- Recorded illustrative cases, persistence across reloads, and a durable export containing the complete replay checkpoint.

The first useful recorded example is a pipe-fault checkpoint at tick 116. Its delivery-at-B test gains 11.46 points over ordinary care. Readings contribute 18.75 points relative to withholding, while the remaining effect is −7.29 points. Both branches initially open B's valve. At tick 128, the branch with readings moves west while the branch without them closes B's valve. This is an illustration selected from the study, not an additional success estimate.

The live controller receives no future state from these comparisons. The old learned selector remains an explicitly optional v0.9 experiment. Maze and image algorithms are unchanged by this stage.

## Cost and verification

The audit (full source archive) verifies 37 frozen scientific source files, earlier source and evidence hashes, complete candidate accounting and the frozen protocol:

| Work | Physical actions |
|---|---:|
| Formal calibration | 4,104 |
| Formal ordinary-care collection | 28,800 |
| Formal paired branches and full-state references | 172,608 |
| **Formal total** | **205,512** |
| Development | 31,404 |
| Development prefix-water instrumentation verification | 1,717 |
| Interactive browser example | 960 |
| Independent export verification, two attempts | 1,920 |

The formal run used **971,631,529 model transitions**, **610.6 seconds active computation**, and **zero uncertain actions**. It stayed within its 205,704-action and 15-minute ceilings. These counters are not FLOPs or energy estimates. Software-test work is separate from the research and replay totals above.

All **83 unit tests** and both HTTP integration suites passed. Browser checks covered synchronized playback, the changed-action jump, reload persistence, durable export and desktop/mobile layouts. The exported comparison reproduced the action sequences, viability outcomes and prediction-step counts exactly in Node. Across 74,827 numerical fields, 106 browser-versus-Node rounding differences remained, with maximum absolute difference 1.43×10⁻¹⁴, below the explicit 10⁻¹⁰ verification tolerance. An initial strict-equality verification attempt failed on those rounding differences and is counted above; no scientific code was changed to erase them. Details are in the verification record (full source archive).

## Next handoff

Keep the working ensemble controller. Do not expand this exposed sample or train another selector merely to finish a version.

The next proposed question is: **Can a shorter test, ended when its reading can change an imminent care decision, retain the information benefit without losing so much time to the test?** Ordinary care already includes local visits; the comparison must show value beyond those existing actions. Use the exposed cases for development, then a separate frozen evaluation with fresh seeds and matched budgets. Compare complete care outcomes, not just the number of readings or changed decisions. This is a hypothesis motivated by the 69 reading-positive but insufficient-total-benefit tests, not a demonstrated fix.

The [roadmap](../ROADMAP.md) remains at Stage 1. Reliable improvement from learned investigation is still unproved.
