# v0.16: a reading forks the plan, but care still loses

The two-decision planner passed its information-boundary checks and sometimes selected different follow-ups after distinguishable local readings. It failed the frozen complete-season development gate: **71.60% both-bed viability versus 82.86% for ensemble care, −11.26 percentage points**. Ten of its fifteen seasons reached the computation reserve limit. The candidate is retired, with no fresh evaluation or live promotion. Ensemble remains the default.

## What changed

The first journey is common to all fitted hypotheses. Along it, the planner predicts only local sensor channels and groups those observation histories into fixed bands. It then evaluates a second ordinary journey within each group. Hypotheses that still predict indistinguishable histories share every forecast action; distinguishable histories can support different follow-ups. The horizon, model support weights, fitting, constant-weather assumption and loss—including terminal reserve valuation—stay fixed.

Only the first journey is committed. Real readings subsequently update the estimator and trigger a new search. A predicted second decision is therefore an expectation, not a binding commitment. At the retained steady plan at tick 115, one group forecasts “Read bed B” after observing A, while another forecasts opening B fully. The next recorded decision at tick 119 is instead “Read shared reservoir.” A two-level tree has not eliminated the inconsistency between its later forecast and actual replanning.

This is a deterministic two-decision approximation inspired by action-observation histories in [Silver and Veness's POMCP paper](https://papers.nips.cc/paper_files/paper/2010/file/edfbe1afcf9246bb0d40eb4d8027d90f-Paper.pdf), not a POMCP implementation. The support weights are not calibrated probabilities; forecast groups depend on authored bin widths and a limited fitted model family.

## Short continuations and complete seasons

The four explanatory cases use exactly the earlier exposed initial journeys and checkpoints. No settings were changed after observing them.

| Exposed 48-tick continuation | Ensemble | Two-decision candidate |
|---|---:|---:|
| Steady, original choice | 56.25% | 87.50% |
| Steady, previously lowest-cost journey | 100.00% | 100.00% |
| Compound, original choice | 100.00% | 91.67% |
| Compound, previously lowest-cost journey | 100.00% | 97.92% |

The candidate improved the original failure but harmed two controls. These selected cases are explanations, not evidence of generalization. None reached its inherited computation allowance.

The full pilot used three exposed seeds, all five conditions, two controllers and 480 ticks per season. Its gate required completion and a nonnegative mean paired viability difference before a fresh comparison. The method and all scientific sources were frozen before execution.

| Seed, mean of five conditions | Candidate | Ensemble | Difference | Candidate seasons limited |
|---|---:|---:|---:|---:|
| 93101 | 45.17% | 50.00% | −4.83 pp | 2 / 5 |
| 93103 | 88.21% | 99.92% | −11.71 pp | 3 / 5 |
| 93113 | 81.42% | 98.67% | −17.25 pp | 5 / 5 |
| All three | **71.60%** | **82.86%** | **−11.26 pp** | **10 / 15** |

These balanced development averages are descriptive. No confidence interval or fresh generalization claim is attached. The ensemble rows exactly reproduce the v0.15 pilot control rows. No ensemble season reached the model limit.

## Computation is a problem, but not the whole explanation

Candidate seasons used 142,949,054 model transitions versus 36,189,496 for ensemble care, approximately 3.95 times as many. Both retained the twelve-million ceiling. The original conservative reservation stops new searches when fewer than two million planning transitions plus 100,000 update transitions remain; “limited” can therefore occur below twelve million. Physical time then continues through waiting actions. It does not mean the entire service or replay stopped.

The retained compound candidate begins this waiting behavior with the action producing state tick 437. Across states 1–436, candidate viability was already **48.39% versus 60.78%** for ensemble care. Across states 437–480, both had zero viability. That case's disadvantage preceded the computation stop. The retained steady candidate used no budget fallback and improved from 44.38% to 52.08%, while remaining poor in absolute terms. These two selected trajectories cannot isolate the causal effect of the budget across the full panel.

The practical inference is to investigate both scoring and search efficiency. Increasing the allowance alone is not supported as a repair. The terminal reserve penalty remains untested as a causal contributor; its large values in v0.15 make it a specific next diagnostic rather than a justification to remove it immediately.

## Costs and verification

| Work | Physical actions | Model transitions | Active time |
|---|---:|---:|---:|
| Four candidate continuations | 192 | 5,873,247 | 4.71 s |
| Full development pilot, including calibration | 15,426 | 179,142,456 | 139.72 s |
| Development total | **15,618** | **185,015,703** | **144.43 s** |

All reservations closed and all pilot rows count. The audit (full source archive) verifies 56 frozen scientific sources, protocol-before-execution ordering, costs, control parity and prior evidence. Independent execution exactly reproduced all four candidate traces (including branching plans) and the four retained complete-season trajectories. That separate verification used 2,454 physical actions and 28,603,013 model transitions.

All 124 automated tests and both HTTP integration suites passed. Delivery verification (full source archive) records exact exports, navigation, playback, budget jumps, responsive layout and the preserved live nursery. The service retains the v0.14 live engine hash; this cycle adds an offline candidate and a read-only evidence view.

**When readings branch** beside the nursery controls opens `/contingencies.html`. Inspect predicted reading groups and their follow-ups, compare the next actual decision, replay short continuations or complete seasons, jump to the first computation-limited action, and export the recorded evidence. Playback performs no simulation.

## Next bounded question

Before another policy, isolate the terminal value and forecast-consistency assumptions on fixed exposed checkpoints. Compare the same candidate actions and physical horizon with the original terminal reserve, no terminal reserve, and an explicit bounded continuation scored by realized care and resource use. Keep the same models, observations and weather assumptions; label any privileged reference. Select cases by a fixed clock or earlier frozen rule, including successful controls. Report rankings and subsequent full-season behavior separately.

Record which calculations are repeated across root candidates and observation groups, with exact accounting. Any reuse must preserve scores and choices; a cheaper tree still needs to pass the full-season gate. Do not retune this rejected candidate or draw fresh cases until the next mechanism and criterion are fixed. Stage 1 remains active.
