# v0.15: the rescue that keeps moving

We found repeated deferral in an exposed nursery failure. A candidate that makes every model evaluate the same future action script repaired that short continuation, but made complete development seasons substantially worse. It failed the predeclared pilot gate and was retired. Ensemble care remains the live default; no fresh evaluation cases were consumed.

## What the trace reveals

The four retained cases use the original choice and the previously lowest-cost initial journey at two v0.14 checkpoints: steady at tick 113 and compound at tick 115, both from seed 93101. They explain behavior in selected exposed cases; they do not measure its prevalence. All 74 recorded planning decisions had at least one difference between the three models' assumed follow-up scripts.

In the failing steady continuation, Lumen repeatedly chose “Read bed A.” At every planning tick from 133 through 146, its selected model predicted closing bed B four ticks later. The expected operation moved from 137 to 150 while actual B moisture rose from 60.66% to 85.67%. It crossed the upper viable threshold of 75% at state tick 141. No B closing operation occurred inside the replay window, which ends at state tick 161. The plan labelled “B valve → 0%” at tick 160 still required travel; that label is not a completed valve operation.

The selected model remained normal and tracked B moisture accurately during this sequence. At tick 133, normal and supply predicted closing at 137, while pipe-1 predicted no closing within its 48-tick horizon. Weighted running cost was 0.208; the estimated terminal reserve penalty was 28.012. That imbalance keeps terminal valuation a competing explanation. The trace identifies a planning inconsistency, but does not isolate it as the sole cause: model support, predicted state and reserve valuation interact.

## One bounded correction, then its rejection

The candidate retains the original candidate menu, three hypotheses and weights, fitting and switching, observation boundary, horizon, weather assumption and cost. For each candidate journey, the selected model generates its heuristic follow-up script, and every hypothesis evaluates those same actions. Only the initial journey is committed; subsequent real readings still produce new decisions. This is a common forecast script, not a policy that branches on future observations.

| Exposed 48-tick continuation | Original ensemble | Common-script candidate |
|---|---:|---:|
| Steady, original choice | 56.25% | 100.00% |
| Steady, previously lowest-cost journey | 100.00% | 100.00% |
| Compound, original choice | 100.00% | 100.00% |
| Compound, previously lowest-cost journey | 100.00% | 100.00% |

Values are the fraction of recorded ticks with both beds viable. The candidate changes the trajectory after the committed prefix, including opening B only halfway earlier in the failing case. Its local success does not show that correcting the later closure alone caused the improvement.

Before running the full-season pilot, its runner recorded a fixed decision: proceed to fresh evaluation only if mean paired viability was nonnegative, without tuning after the pilot. The pilot used three already exposed seeds, all five conditions, both policies and 480 ticks per season: 30 complete seasons with shared calibration.

| Exposed seed, mean of five conditions | Common script | Ensemble | Difference |
|---|---:|---:|---:|
| 93101 | 43.71% | 50.00% | −6.29 pp |
| 93103 | 78.38% | 99.92% | −21.54 pp |
| 93113 | 78.04% | 98.67% | −20.63 pp |
| All three | **66.71%** | **82.86%** | **−16.15 pp** |

All three seed means were negative; no season exhausted its computation allowance. These are descriptive development results, without a fresh generalization estimate or confidence interval. The candidate failed the gate. Its implementation and completed evidence are frozen, and it is not offered as a live controller.

## Cost and verification

| Work | Physical actions | Model transitions | Active time |
|---|---:|---:|---:|
| Original plan traces, including forecast instrumentation | 192 | 2,690,276 | 1.99 s |
| Candidate continuations | 192 | 315,359 | 0.39 s |
| Complete-season development pilot, including calibration | 15,426 | 48,894,006 | 31.66 s |
| Total development | **15,810** | **51,899,641** | **34.04 s** |

All recorded reservations closed. Forecast instrumentation runs on clones and is charged separately from care; it cannot affect the original controller or its allowance. The audit (full source archive) verifies all 54 frozen scientific sources, retained records and prior evidence. The live engine retains its v0.14 hash because this cycle changes no live policy or worker dispatch.

Independent execution exactly reproduced all four original/candidate trace pairs, including actions, frames, recorded plans and metrics, plus calibration and four complete-season metric records. That verification separately cost 2,646 physical actions and 8,530,121 model transitions. The 119-test suite and both HTTP integration suites passed. Delivery verification (full source archive) records browser controls, exports, responsive checks and preserved nursery state.

**Follow the plans** beside the nursery controls opens `/plans.html`. It shows expected operation dates against planning dates, local readings, model support, costs, actual operation timestamps, paired recorded continuations and the failed pilot. Playback adds no simulation. A missing operation is bounded by the recorded horizon, not a claim that it never occurs.

## Where to go next

Keep the ensemble and test one small observation-conditioned planning tree. Models that predict indistinguishable observable histories must share actions; a later local reading may justify a different follow-up. Begin with two decisions and explicit observation grouping. Test that indistinguishable histories cannot secretly select different actions, and charge every simulated branch. Keep terminal reserve valuation fixed initially so that the change is interpretable. If this fails the exposed full-season gate, reconsider terminal valuation rather than tuning repeatedly on fresh cases.

The relevant primary reference is [Silver and Veness (2010), *Monte-Carlo Planning in Large POMDPs*](https://papers.nips.cc/paper_files/paper/2010/file/edfbe1afcf9246bb0d40eb4d8027d90f-Paper.pdf). POMCP organizes search by action-observation histories and combines simulation with belief updating. The proposed two-decision adaptation borrows that structure; it is not a reproduction of POMCP. Our heuristic weights are not calibrated beliefs, and fitted dynamics and observation grouping require their own checks. The paper's guarantees do not automatically apply here.

Stage 1 remains active: reliable improvement from experience is still unproved. The next candidate earns fresh evaluation only after a frozen development gate; this cycle's local repair is not a reason to skip that gate.
