# v0.17: terminal value and repeated forecast work

Freeze this method, implementation and runners before observing diagnostic results. Preserve the prior 56 scientific sources and all earlier evidence. The live ensemble and its engine hash remain unchanged.

## Fixed exposed panel

Reuse the v0.14 panel: six exposed base seeds 93101, 93103, 93113, 93131, 93133 and 93139; scenario offset 9,100,001; steady, pipe, dry, supply and compound; 480 ticks. Collect the first idle planning point at or after ticks 112 and 304, with at least 96 remaining ticks. Calibration and collection are new charged executions on exposed cases. Keep the first seed's first steady and compound checkpoints as explanatory examples regardless of outcomes. This yields 60 checkpoints and includes successful controls.

## Score the same twelve initial journeys

Use identical fitted hypotheses, weights, initial state estimates, candidate order, current weather held constant and running loss. For each initial journey, evaluate the frozen model-specific heuristic forecast. Compare:

1. **Reserve:** the original 48-tick running cost plus its terminal water-reserve penalty for the remaining season.
2. **None:** the identical first 48-tick running cost without that penalty.
3. **Extension:** the identical first 48 ticks followed by 48 more explicitly simulated heuristic ticks, with their care, water, travel and drain costs; no final reserve penalty.

The first 48 actions and frames must exactly match between the short and extended predictions. The extension is a bounded alternative terminal evaluation, not an estimate of the entire rest of the season. Its extra model work counts. Original scores must match the frozen planner. No actual hidden state, future disturbance coefficients or future weather enters scoring.

For every journey, independently execute two 96-tick physical continuations from the same private checkpoint: (a) commit only its initial prefix, then resume unchanged ensemble care; (b) execute its selected-model 96-tick forecast script. The latter is an offline observer intervention, not a capability supplied to Lumen. Score actual care loss and water use at ticks 48 and 96, both-bed viability, and first action divergence. These branches distinguish a forecast-script assumption from the consequences of later replanning. Models can still predict incorrectly in both branches.

Primary descriptive ranking outcome: mean actual 96-tick regret relative to the lowest realized-cost candidate at each checkpoint. Also report 48/96-tick viability, water, exact or tied best selections, choice changes, and seed means. These are retrospective diagnostic opportunities, not achieved policy gains or fresh inference. No confidence interval is assigned to exposed results.

Physical ceiling: 6 × 342 calibration + 30 × (480 collection + 2 × 12 × 2 × 96 branches) = **154,692 actions**. Reserve 342 for calibration and 5,088 per habitat before execution. Maximum 1.5 billion model transitions, four million per care branch, and 600 active seconds; check ceilings between bounded habitat chunks. All models, rejected candidates, fitting and diagnostic repetitions count. Retain uncertain reservations on failures instead of claiming completion.

## Conditional nomination

Nominate one scoring ablation for a separately frozen exposed full-season pilot only if it lowers mean 96-tick regret, raises mean 96-tick viability, and has nonnegative viability differences on at least four of the six seed means. Prefer higher viability, then lower regret. A nomination is not promotion. If neither qualifies, retain the diagnostic result without creating another tuned controller. Any pilot must be fixed before execution; any fresh evaluation requires a later independent gate.

## Repeated-work measurement

Profile one complete two-decision search at each of the two previously retained v0.14 checkpoints. Use instrumented copies of the frozen forecast and tree routines. Every model transition still executes. For each call to the physical model, key the full relevant fitted state (tick, position, tank, weather, coefficients, moisture, equipment and pipe queues) plus action. Count total and distinct keys, repeated transitions, and repeats across root journeys. This is an exact reuse opportunity, not achieved acceleration or a runtime guarantee; lookup, storage and serialization have their own costs.

Compare every instrumented root evaluation byte-for-byte with the frozen v0.16 result and charge both executions. Keep projection work separate from rollout work. Profile ceiling: two million total transitions including parity execution and 60 active seconds; zero physical actions. Do not implement cache reuse or change the candidate's budget behavior in this cycle solely because repeated keys exist.

## Verification and delivery

Before running, use software fixture seed 90101 to verify core-prefix and score parity, exclusion of hidden future changes, exact extension of the original 48-tick physical continuation, unchanged source checkpoints, physical/model counts, instrumented tree parity, and seed-consistency nomination rules. After running, audit all hashes, protocol ordering, rows and cost totals; replay retained examples independently. Add a view of score components, competing rankings and paired script/replanning trajectories, with exact exports. Keep recorded evidence, development computation and software verification separate.
