# v0.11: When to stop investigating

Completed on 19 September 2026 (America/New_York). Timely care averaged 84.92% both-bed viability, versus 83.91% for ensemble care. The paired gain was +1.01 pp (95% seed-cluster interval -0.14 pp to +2.17 pp). The timely policy did not meet the frozen promotion rule. Ensemble care remains the default.

## What changed in the nursery

**See timed tests**, beside the environment controls, opens five-policy results and three synchronized recorded worlds. Select short tests, long tests or the existing investigator alongside timely and ordinary care. Play, scrub, jump to the first early stop, and export all five recorded journeys with their frozen provenance. Recorded playback performs no new simulated actions.

The new optional next-season controllers are timely, short and long tests. Their live status reports test duration and the choices that ended a wait. A complete timely season survives reload and exports the initial calibration, configuration, action trace, test history and model cost. The previous reading-withholding comparisons, maze links and older evidence remain available.

In the first recorded early-stop example (steady conditions, base seed 93101), at tick 142 the bed B reading is about 78%. With the test's new readings, the next care choice is **Drain B**; the shadow learner without them would choose **B valve → 100%**. Lumen ends the test after five of six planned ticks, after completing its route and operation. This is an explanation of an actual recorded choice; it does not demonstrate a correct diagnosis or make this selected example representative.

## The complete comparison

Twelve fresh base seeds × five conditions × five policies = **300 full 480-tick seasons**. Each policy gets the same action and model-transition ceilings. Calibration is shared within a seed; learning memory is reset between condition/arm pairs. The primary estimate clusters the five conditions within each independent seed.

| Policy | Both beds viable | Total test ticks | Early stops | Model transitions | Computation-limited seasons |
|---|---:|---:|---:|---:|---:|
| Timely | 84.92% | 956 | 76 | 363,554,145 | 0/60 |
| Short fixed duration | 84.43% | 1,243 | 0 | 355,498,259 | 0/60 |
| Long fixed duration | 83.89% | 1,243 | 0 | 207,511,646 | 0/60 |
| Ensemble | 83.91% | 0 | 0 | 153,559,776 | 0/60 |
| Existing value tests | 83.64% | 999 | 0 | 206,835,553 | 0/60 |

The promotion rule required a mean primary gain of at least one percentage point and a lower 95% interval above zero. The timely policy did not meet the frozen promotion rule. Ensemble care remains the default. This was a superiority check; it does not establish equivalence or noninferiority when that threshold fails.

| Paired contrast | Mean | 95% interval |
|---|---:|---:|
| timely-ensemble (primary) | +1.01 pp | -0.14 pp to +2.17 pp |
| timely-short (exploratory) | +0.49 pp | -0.14 pp to +1.11 pp |
| short-long (exploratory) | +0.54 pp | -0.62 pp to +1.71 pp |
| timely-value (exploratory) | +1.28 pp | -0.22 pp to +2.79 pp |

Timely care ended 76 tests early and saved 211 ticks against those tests' scheduled durations. It used 956 test ticks across all seasons, versus 1243 for short fixed-duration care. These policies can select different later tests and visit different states. The whole-policy contrasts do not isolate the causal effect of one reading or of removing a particular wait. Secondary intervals have no multiplicity adjustment.

| Condition | Timely | Short | Long | Ensemble | Existing value |
|---|---:|---:|---:|---:|---:|
| steady | 90.71% | 90.68% | 89.34% | 89.50% | 89.50% |
| pipe | 89.06% | 87.08% | 87.03% | 86.84% | 86.72% |
| dry | 80.05% | 80.07% | 79.95% | 79.98% | 79.08% |
| supply | 82.93% | 83.06% | 81.65% | 81.58% | 81.67% |
| compound | 81.82% | 81.27% | 81.48% | 81.63% | 81.22% |

Condition means are descriptive. The decision uses the preregistered aggregate, not whichever condition looks strongest.

## Cost and verification

The service charged **148,104 physical actions**: 4,104 calibration and 144,000 care actions; zero uncertain actions. All model fitting, forecasting, shadow updates and stopping previews totaled **1,286,975,003 model transitions**. Active worker time was 862.5 seconds, within the 1,200-second ceiling. Model transitions are not a measurement of energy or total application CPU cost.

Both complete development runs are preserved separately, each with 24,684 physical actions. The unscreened run exhausted its planning allowance and harmed care. The screened development run stopped exhausting its budget but still did not improve care. The intermediate end-of-season error retains 3,222 completed actions plus a 480-action interrupted reserve; its accounting review adds the calibration model cost without changing the frozen record. See the development accounting (full source archive).

All **40 frozen scientific source files** match their manifest; older scientific sources and report hashes are unchanged. **92 unit tests** and both HTTP integration suites passed. Five exported recorded seasons reproduced exactly in Node: 2,400 separately counted physical replay actions and 16,309,327 model transitions. The live browser season uses 480 actions; two saved-checkpoint reloads reconstruct another 960. A further independent Node replay reproduced the 480 browser actions and cost counters exactly, with only floating-point differences below 2.7e-15 in numeric trace fields. None of these checks adds validation cases. Desktop and 390-pixel layouts, comparison switching, playback, the no-early-stop case, exports and a paused ensemble restart were checked in the browser.

Artifacts: [frozen method](v11-method.md), manifest (full source archive), full report (full source archive), audit (full source archive), recorded replay verification (full source archive), browser journey verification (full source archive), and software verification (full source archive).

## Where this leaves the roadmap

The implementation now demonstrates decision-dependent stopping. Its existence and saved waiting time have not met the declared criterion for better care. The models, test vocabulary, screening and stopping rules are authored. This is not a learned investigation selector or evidence of general intelligence or sentience.

Stop tuning test duration for now. Diagnose why predicted care choices fail: compare the same exposed checkpoints under fitted versus known dynamics, and the existing versus a longer planning horizon, with matched computation budgets. Use that factorial diagnostic to distinguish model error from planning limits before choosing another policy change. The diagnosis is proposed work, not a conclusion that either component is already known to be at fault.

Keep the [roadmap](../ROADMAP.md) grounded in this distinction between visible behavior and measured benefit. More test starts or a larger archive does not establish cumulative learning.
