# v0.14: selected-model scoring in fresh complete seasons

## Development decision

The fixed exposed journey audit completed 60 checkpoints and 1,440 matched branches. Existing weighted scoring selected a lowest-cost ordinary continuation at 25/60 checkpoints, with mean running-cost regret 4.2067. Scoring with the active model alone selected one at 22/60, but reduced mean regret to 1.6475. The latter changed checkpoint-window both-bed viability from 87.6042% to 88.9236% (+1.3194 percentage points). Most of that benefit came from base seed 93101; it was slightly worse on 93139 and unchanged in viability on the other four seeds. This is a narrow development signal, not generalization.

All tested continuations eventually differed from the assumed script. Exact physical script outcomes still selected the lowest-cost ordinary continuation at only 21/60 checkpoints. Prediction errors and tail mismatch interact: mean selected-model cost error was actually smaller against ordinary care (19.6780) than against the physical script (22.7530). The audit does not support treating a tail rewrite as an isolated, proven fix. Giving the weighted planner future weather also failed to improve exposed choices overall.

One small candidate follows from the severity of the weighting mistakes: retain the entire hypothesis bank and its original observation-based fitting, validation and switching, but evaluate each ordinary journey with the currently active model only. Do not tune thresholds, supports, model families or tail settings. This isolates the decision score. It also saves forecast computation, but care improvement is the sole promotion claim.

## Candidate and boundaries

The candidate receives the same local observations as ensemble care. The action menu, order and stable tie breaking, 48-tick horizon, constant current-weather assumption, heuristic follow-up, running loss and terminal reserve penalty are unchanged. Model switching can interrupt a pending journey just as in live ensemble care; unlike the exposed audit, this evaluation does not force any initial prefix. All hypotheses continue receiving observations, fitting and validation. The active model's forecast cost receives weight one; no other hypothesis contributes to the decision score.

Software verification uses already exposed seed 90101. The selected planner must reproduce the audit's selected-member scores and action prefixes, and its ensemble comparison delegates directly to the frozen controller. No fresh evaluation cases are used for implementation or tuning.

## Frozen evaluation

Fresh base seeds: 124101, 124103, 124107, 124109, 124121, 124123, 124133, 124139, 124147, 124153, 124159, 124171. Scenario seed is base plus 13,100,001. Each has steady, pipe, dry, supply and compound conditions and two arms: selected and ensemble. Every season completes 480 ticks from the initial habitat; the two arms share the unchanged 342-action calibration within a seed.

Primary outcome: selected minus ensemble complete-season both-bed viability, averaged over five conditions within a seed, then across twelve independent base seeds. Use a paired 95% t interval with critical value 2.201. Promote only when mean gain is at least one percentage point and the interval's lower endpoint is above zero. A partial or budget-ended study cannot promote. Model transitions, water and active worker time are descriptive secondary outcomes; do not substitute a noninferiority or efficiency claim after seeing the results.

Budget: twelve chunks of 5,142 physical actions, totaling **61,704** including calibration. Each season has the original twelve-million model-transition ceiling, with two-million planning and 100,000 update reservations. Limited seasons continue with observed waiting and remain in the analysis. The service reserves physical actions durably and has a fifteen-minute active-worker ceiling, including interrupted work. All actual hypothesis updates and rejected candidate forecasts count. This fresh comparison's budget is separate from the 85,572 physical actions and 261,583,873 model transitions used in development.

Retain both arms' complete first-seed steady and compound seasons for synchronized replay regardless of result. Freeze the implementation, this protocol, diagnostic report and earlier scientific source hashes before starting the first fresh seed. Preserve all prior reports and the current saved nursery journey. If the promotion rule fails, keep ensemble care as the default and use these newly exposed cases only for subsequent diagnosis.
