# v0.16: one observation-conditioned second decision

This protocol is written and frozen before the development continuations or full-season pilot. The candidate is offline only. It earns a fresh evaluation only by passing the gate below; development success cannot authorize a live promotion.

## Candidate and information boundary

Retain the existing three fitted hypotheses and heuristic support weights, model fitting and switching, ordinary twelve-journey menu and stable tie order, 48-tick horizon, constant current weather, running loss and terminal reserve penalty. Do not change the reserve penalty in this cycle.

For each initial journey, simulate its complete prefix in every fitted hypothesis, collecting only locally available sensor channels after each action: position and tick, visible bed moisture, nearby equipment, and tank level only within sensing range. Quantize moisture into fixed 0.02-wide bins and tank level into 0.05-wide bins, anchored at zero using floor. These bins define a computational approximation; they are not a claim about measurement noise. Weather is held constant and common to all hypotheses, so cannot distinguish branches.

Group hypotheses with identical quantized observation histories along that journey. Within each group, enumerate the ordinary menu again. The current selected model supplies the heuristic script if present; otherwise use the highest-support member, with stable original order for ties. Every member in the group evaluates the same complete child script. Choose the child with the lowest support-weighted cost. Combine the initial prefix cost with the groups' weighted future costs, counting the terminal reserve once. All models share initial actions. After that journey, different forecast actions require distinguishable local histories. The forecast has exactly two explicit menu decisions; the remaining tail is a common script within each branch.

Only the initial journey is executed. Subsequent real readings feed the unchanged estimator and a new search. Predicted branch membership is not a claim that the actual model is known, and a forecast child is not a commitment to execute it. No actual hidden state, disturbance schedule or future weather is supplied to planning.

This borrows the action-observation-history principle from [Silver and Veness (2010), Monte-Carlo Planning in Large POMDPs](https://papers.nips.cc/paper_files/paper/2010/file/edfbe1afcf9246bb0d40eb4d8027d90f-Paper.pdf). It is a deterministic two-decision approximation, not a POMCP implementation or calibrated Bayesian planner. That paper's convergence guarantees are not transferred to this system.

## Checks before evidence

Using software fixture seed 90101, verify the forecast sensor mask against real local observation channels; indistinguishable histories sharing all future actions despite different hidden predictions; invariance to model renaming; different readings permitting different follow-ups only after the shared prefix; exact physical-state parity with the frozen simulator, running-cost parity within floating-point tolerance and one terminal reserve; accounting for every rejected child; end-of-season behavior; exact ensemble-control delegation; and the original model ceiling.

## Exposed explanatory continuations

Use the same four retained v0.15 cases: original and previously lowest-cost prefixes at the v0.14 steady tick-113 and compound tick-115 checkpoints, base seed 93101. Run the candidate for 48 ticks after committing each identical prefix. Compare with saved original traces. These selected cases explain behavior; they neither select candidate settings nor estimate prevalence or authorize fresh evaluation. Reserve 192 physical actions, 48 million model transitions and 120 active seconds. Keep inherited controller compute usage; limited continuations count.

## Fixed full-season pilot and gate

Use already exposed seeds 93101, 93103 and 93113, scenario offset 9,100,001, all five conditions (steady, pipe, dry, supply, compound), both contingent and ensemble controllers, and 480 ticks per season. Share the unchanged 342-action calibration within each seed. No forced prefix in complete seasons. Retain the first seed's steady and compound paired trajectories for explanation regardless of outcome.

Physical ceiling: 3 × (342 + 5 × 2 × 480) = 15,426 actions. Reserve 5,142 per seed before its work. Both policies keep the original 12-million per-season model ceiling with 2-million planning and 100,000 update reserves. Charge every prefix, rejected child, tail projection and model update. Pilot ceiling: 360 million model transitions and 300 active seconds, checked between bounded seed batches. A ceiling breach stops the pilot and blocks fresh evaluation; do not present a partial result as completion.

Gate: proceed to one frozen fresh comparison only if the pilot completes and mean paired both-bed viability is nonnegative. Average five conditions within each seed, then three seed means. Limited seasons count in the result. Do not tune the implementation or bin widths after seeing either continuation or pilot outcomes. A negative result retires this candidate. These exposed results are descriptive, without confidence intervals or a fresh generalization claim. If the candidate passes, write and freeze the fresh protocol before choosing new seeds or running evaluation; ensemble remains the live default until that independent gate passes.

## Preservation and delivery

Preserve all 54 prior frozen scientific sources and earlier evidence. Add two scientific modules and a separate analysis manifest; the live service retains its v0.14 engine hash. Record completed costs and reservations. Add an inspectable observation-branch view and paired recorded replays, clearly distinguishing forecast branches from actual replanning. Verify exact replay, exports, service boundaries, responsive layout and preserved live nursery state. Playback costs no new simulation; verification costs are separate from development.
