# v0.17: a large reserve penalty does not explain these choices

Removing the terminal water-reserve penalty changed **0 of 60 initial choices** on the fixed exposed panel. Extending the forecast from 48 to 96 ticks changed **35 of 60**, but reduced mean 96-tick both-bed viability from **87.03% to 86.96% (−0.07 percentage points)**. Neither scoring change meets the frozen nomination rule. No full-season candidate or fresh evaluation follows from this diagnostic. Ensemble remains the live default.

## What was held fixed

Six previously exposed seeds, five conditions, and two checkpoints per season yielded sixty comparisons. Every checkpoint retained all twelve authored initial journeys in their original order, the same local observations, fitted models and support weights, and the same constant-weather assumption. Every journey was scored three ways:

- The original 48-tick running cost plus estimated water-reserve penalty for the rest of the season.
- Exactly the same 48 predicted ticks without that penalty.
- Exactly those 48 ticks followed by 48 more explicit heuristic forecast ticks, without a final penalty.

The last alternative changes the terminal assumption and the explicitly simulated horizon. It does not simulate the whole remaining season. Each hypothetical model still supplies its own heuristic follow-up; this diagnostic does not solve the forecast/replanning inconsistency or replace it with the rejected v0.16 tree.

After each identical initial journey, an offline branch resumed unchanged ensemble care for the remainder of a 96-tick window. A separate branch executed the selected fitted model's complete forecast script for that window. Actual hidden state and future weather never entered scoring. The observer script is an experimental intervention, not a capability available to the live controller.

## Ranking and care

| Score | Changed choices | Both-bed viability, 96 ticks | Mean realized regret | Best realized continuation selected | Mean water |
|---|---:|---:|---:|---:|---:|
| Original 48 + reserve | — | 87.0313% | 6.7418 | 17 / 60 | 0.9664 |
| 48, no reserve | 0 / 60 | 87.0313% | 6.7418 | 17 / 60 | 0.9664 |
| 96, no reserve | 35 / 60 | 86.9618% | 6.3949 | 9 / 60 | 0.9796 |

Regret is actual cost above the best of the twelve recorded ensemble continuations at the same checkpoint. The extension reduces mean regret somewhat while slightly reducing viability and choosing a cost-optimal continuation less often. These outcomes are compatible: the measures value different aspects of care and error severity. None is a fresh generalization estimate or a complete-season policy result.

| Exposed seed, mean of ten checkpoints | Original / no reserve | Extension | Difference |
|---|---:|---:|---:|
| 93101 | 45.3125% | 45.1042% | −0.2083 pp |
| 93103 | 100.0000% | 99.4792% | −0.5208 pp |
| 93113 | 99.5833% | 99.5833% | 0.0000 pp |
| 93131 | 100.0000% | 100.0000% | 0.0000 pp |
| 93133 | 99.1667% | 99.3750% | +0.2083 pp |
| 93139 | 78.1250% | 78.2292% | +0.1042 pp |

The frozen rule required lower mean regret, higher mean viability, and nonnegative viability differences on at least four of six seed means. The extension meets the last criterion but fails the viability requirement. Removing reserve changes no choices. Thus **neither earns even the separately frozen development pilot**; this result does not justify changing the live controller.

The large absolute reserve scores seen earlier were a plausible lead, but large values need not change candidate ordering. In the retained steady checkpoint, several competing journeys have virtually the same 7.8502 reserve penalty. Removing it leaves the original choice, “A valve → 50%,” preferred. Reserve differs for other journeys, so this is not a claim that it is always constant or irrelevant in other states or planners.

## The promise still differs from the actual continuation

Every one of the 720 selected-model scripts eventually differed from actual ensemble replanning. That count alone does not establish harm: real readings can appropriately change a plan. The paired branches make consequences inspectable.

At the retained steady checkpoint, tick 113, the original choice's script achieves **97.92%** both-bed viability over 96 ticks. Committing the same prefix and then replanning achieves **69.79%**, with the first action difference producing tick 116. All three score rules still choose that initial journey. The compound checkpoint at tick 115 provides a successful control: the original “Read bed A” choice achieves **100%** in both branches. These examples were fixed before this run's outcomes.

This narrows the diagnosis: terminal penalty removal does not repair the choices tested here. It does not prove that blindly committing to forecasts is a safe controller. A fixed script can fail elsewhere, and model error and future weather remain part of the discrepancy.

## Repeated work is measurable

Two full searches from the rejected v0.16 two-decision candidate were instrumented separately. Every physical-model call ran; the profiler counted repeated exact fitted-state/action keys and compared each full root result with the frozen original.

| Search | Instrumented transitions | Distinct keys | Repeated calls | Repeated across root journeys |
|---|---:|---:|---:|---:|
| Steady, tick 113 | 70,215 | 22,727 | 47,488 (67.63%) | 7,673 |
| Compound, tick 115 | 115,199 | 24,990 | 90,209 (78.31%) | 10,720 |

Most repeated calls occurred in short lookahead projections: 41,901 of 51,195 projection calls in steady, and 81,817 of 96,047 in compound. This suggests starting reuse there. The profile did **not** implement a cache, reduce model work, or measure a speedup. State serialization, lookup, storage and copying can offset savings. Two exposed searches cannot establish a general performance result.

## Costs and verification

| Work | Physical actions | Model transitions | Active time |
|---|---:|---:|---:|
| Calibration, collection and all diagnostic branches | 154,692 | 462,159,871 | 248.06 s |
| Search profile, including original-result parity | 0 | 370,828 | 0.93 s |
| Development total | **154,692** | **462,530,699** | **248.99 s** |

All reservations closed. All 60 checkpoints and 720 pairs of 96-tick continuations completed, and no branch reached its model ceiling. Counts include fitting, rejected candidates, lookahead projections and parity runs. Profile runtime includes instrumentation and its comparison execution; it is not an ordinary planner benchmark.

The audit (full source archive) verifies 58 frozen scientific sources, method-before-run ordering, all earlier evidence, costs, complete panel coverage, and exact parity with all **720 earlier 48-tick comparisons**, including both physical branches. Independent replay reproduced all 24 retained journeys, their forecasts and both 96-tick branches exactly. That separate verification used **4,608 physical actions and 14,313,842 model transitions**, taking 7.75 active seconds. Software fixture tests and UI checks are separate from development evidence.

The new **What is later worth?** link beside the nursery controls opens `/reserve.html`. Inspect score components, each model's water estimate, all twelve rankings, paired script/replanning replays, seed means, and repeated-work counts. Exports contain the exact recorded comparison; playback performs no simulation. The live engine hash remains the v0.14 hash.

All **128 automated tests** and both HTTP integration suites passed. Delivery verification (full source archive) records exact durable exports, desktop and mobile checks, working playback and end restart, scoring controls, expandable evidence and nursery navigation. The saved live nursery remains paused at tick 230/480 with ensemble care.

## Next bounded work

First, test a bounded exact cache for repeated fitted-model transitions, beginning with lookahead projections. Freeze its scope and eviction rule; require exact score, action and trajectory parity, and count requests, cache hits and actual model executions separately. Benchmark wall time and memory as well as computation on fixed development cases. Preserve the existing budget behavior for a parity check before testing any cheaper-budget controller. A faster implementation of a rejected policy is still a rejected policy until a separately declared full-season gate supports it.

Separately, inspect the first consequential replanning divergence in fixed exposed failures and successful controls: which new reading, belief update or changed follow-up reverses the useful forecast? Use paired interventions to isolate that decision before training another selector. Avoid another round of weight or horizon tuning without an identified mechanism.

Stage 1 remains active. We have ruled out one specific proposed repair on this panel and identified a concrete computation opportunity; reliable improvement from accumulated experience remains unproved.
