# v0.15: tracing deferral and testing a common follow-up script

This cycle contains exposed diagnosis and a fixed development pilot. It does not contain a fresh evaluation. No controller is promoted.

## Trace panel

Use the two retained v0.14 checkpoints: base seed 93101, steady at tick 113 and compound at tick 115. For each, replay the original chosen initial journey and the previously lowest-cost initial journey. These four explanatory cases are selected from already exposed results, including a clear failure and successful controls. They cannot estimate failure prevalence.

Each continuation commits its initial prefix and then runs the unchanged ensemble controller for 48 ticks. Before every subsequent new plan, an instrumented clone records the ordinary candidate scores, all three models' forecast action sequences, local observations, model identities and support, estimated moisture, running cost and terminal reserve penalty. Instrumentation cannot modify the controller's memory or spend its runtime allowance. Its model transitions are added separately to the total research cost. Actual actions and care transition counts must exactly match the retained v0.14 continuations.

Action timestamps identify the state before the action. A closing operation at tick 137 would produce the state at tick 138. The selected-model forecast's first future matching operation is shown against the tick when its plan was created. Missing operations are censored by the 48-tick forecast horizon, not claimed to be absent forever. A changing forecast date alone is not a causal proof of harm.

The trace protocol reserves 192 physical actions, twelve million total model transitions and 60 seconds. It records source hashes before execution and writes progress after each continuation. The protocol and case-selection rule are contained in `src/plan-trace.mjs` and its runner. This document describes that recorded protocol; it is not a retrospectively claimed preregistration.

## One candidate

The trace shows forecasts in which different models assume different later actions despite having the same initial journey. The candidate retains the original three hypotheses and weights, observation boundary, fitting, switching, ordinary menu and stable menu tie order, 48-tick horizon, constant-weather assumption, cost and terminal reserve penalty.

For each candidate prefix, the currently selected model generates the existing heuristic follow-up script. Every other model then evaluates that identical full action sequence. Score the weighted sum of their distinct predicted physical costs. Only the initial prefix is executed; later real observations still trigger ordinary new decisions. No future weather, disturbance schedule, hidden physical state or model identity from the real world enters the policy. This common open-loop forecast script is an approximation, not a contingent policy that can react to future readings.

Verify that every model receives identical actions; its resulting state and cost match the frozen physical-model simulator. When the hypotheses have identical states and coefficients, candidate costs and choices must match the original planner. The ensemble control arm delegates directly to the frozen controller. Test with exposed seed 90101.

Replay the same four initial journeys with the candidate for another 192 physical actions. These comparisons explain its local behavior; they do not select its settings or authorize promotion. Reserve twelve million transitions and 60 seconds. Record source hashes before running and retain the initial reservation until the full batch completes.

## Fixed complete-season pilot

Use exposed base seeds 93101, 93103 and 93113, scenario offset 9,100,001, and all five conditions: steady, pipe, dry, supply and compound. Compare common-script and ensemble care for 480 ticks from each initial habitat. Share the unchanged 342-action calibration within each seed. No prefix is forced in these complete seasons.

Physical ceiling: 3 × (342 + 5 × 2 × 480) = **15,426 actions**. Reserve 5,142 per seed. Retain the original twelve-million per-season model ceiling and conservative planning/update reservations. Total pilot ceiling: 360 million model transitions and five active minutes, checked between bounded seed chunks. Limited seasons count in the results. Counts include calibration, all hypotheses' updates and every rejected candidate forecast.

The pilot runner records its source hashes and decision rule before executing any seed: **proceed to one fresh full-season comparison only if mean paired both-bed viability is nonnegative; do not tune the implementation after this pilot.** Average five conditions within each seed, then three seed averages (equivalent to the balanced 15-pair mean). These exposed results are descriptive. No confidence interval or fresh generalization claim is assigned to them. A negative result retires this candidate without consuming fresh cases.

## Preservation and delivery

All prior 51 scientific sources and frozen records remain unchanged. Freeze the three new scientific files, runners, method and all completed records under a separate analysis manifest. The live service retains the v0.14 engine hash because this cycle changes no live controller or experiment dispatch. Add a read-only planning trace page, visible model disagreements, synchronized original/candidate playback and exact exports. Playback adds no physical or model transitions. Verify source hashes, complete accounting, original trajectory parity, candidate replay parity, service boundaries and the preserved live nursery state.
