# v0.14: does the planner rank care journeys correctly?

This is a bounded development diagnostic on already exposed habitats. It cannot promote a controller or establish fresh generalization. The live ensemble controller and all 47 v0.13 scientific sources remain unchanged.

## Fixed panel

Six exposed base seeds: 93101, 93103, 93113, 93131, 93133, 93139. Scenario seed is base plus 9,100,001. Each uses steady, pipe, dry, supply and compound conditions for 480 ticks. Share the unchanged 342-action calibration within each seed. Collect the first ordinary plan boundary at or after ticks 112 and 304, after assimilation of the current local observation. The two checkpoints must leave at least 48 ticks. These are clock selections, not selections based on failure or hidden disturbance labels.

At each of 60 checkpoints, record the existing 12 ordinary journey prefixes and their three-model weighted forecast costs before executing counterfactual outcomes. Also record scores from the selected model alone and from an offline reference supplying the actual next weather sequence to the same three models. The latter retains fitted states and coefficients and the existing heuristic tail. Its future weather is unavailable to live care. All three scores retain the existing terminal reserve penalty; weights and parameters are not tuned.

## Matched continuations

For every prefix, clone the checkpoint twice and execute 48 physical ticks:

1. **Forecast script:** execute the selected model's complete predicted action script, including its heuristic follow-up, on the real simulator. It receives no corrective replanning. This tests the physical outcome of that assumed script; it is not an optimal or implementable foresight policy.
2. **Ordinary continuation:** force the complete initial prefix, then return to ordinary ensemble care using only its local observations, existing model fitting, switching and planning. Memory starts from the same checkpoint. No hidden world values enter its decisions.

Both branches have identical initial states and exogenous weather. The initial prefix is committed in both, even if a model switch would have interrupted it in unrestricted live care. Therefore this is a comparison of committed journeys, not a replay of the unrestricted baseline controller. All sensing, travel and operating ticks count. Candidate outcomes are dependent counterfactuals, not independent samples.

The instrumented forecast must exactly reproduce the frozen simulator's costs, states and transition counts. Record its chosen actions so that forecast script versus ordinary continuation can be replayed directly. Retain every candidate's numerical result. Keep full visual examples for the first seed's first steady and compound checkpoints, regardless of outcome.

## Outcomes and interpretation

The diagnostic ranks realized **48-tick running care cost**: the existing per-bed viability and margin loss, plus existing water, travel and drain charges. It excludes the predicted terminal reserve penalty because that term is a model estimate of later need, not an observed 48-tick outcome. Record total predicted cost and predicted running cost separately. Thus ranking error can also reflect terminal valuation and the finite observation window. Both-bed viability and the retrospective maximum viable fraction are reported separately; a minimum-cost option need not maximize viability.

For each score, report the fraction of checkpoints at which its first-ranked option ties the lowest realized cost (tolerance 1e-9), mean cost regret and realized viability. Report per-seed means, not confidence intervals that treat 60 checkpoints or 720 candidate journeys as independent. These exposed six-seed summaries are descriptive. No hypothesis test or promotion threshold is applied.

To investigate the tail, compare selected-model predicted running cost with the physical script outcome, then with ordinary continuation. Report absolute error, signed and absolute changes in cost when replanning replaces the script, and the frequency of different action sequences. Also rank by the physically executed script costs and evaluate those choices under ordinary continuation. Even exact knowledge of script outcomes may rank ordinary continuations incorrectly. This is an order-dependent counterfactual comparison; errors and improvements are not an additive causal decomposition.

Retrospective best-candidate gains are headroom within this menu and window. They are not achieved policy gains, optimal-control bounds or evidence that learning a selector will capture the headroom. A future correction must specify how it uses only available observations, then be frozen before a fresh complete-season comparison. Do not choose fresh seeds during this diagnostic.

## Budgets and records

Maximum physical actions: 6 × 342 calibration + 30 × 480 collection + 60 × 12 × 2 × 48 branches = **85,572**. Each habitat reserves 2,784 physical actions before running; calibration reserves 342. Interrupted reservations remain visible and are conservatively charged as uncertain. No retries silently erase exposure.

All collection, calibration, candidate scoring, duplicate parity forecasts, weather-reference forecasts and ordinary branch model transitions count, including rejected candidates. Script execution counts physical actions but performs no model updates. A branch may use at most two million model transitions; the unchanged controller also retains its original twelve-million season ceiling including its inherited prefix history. Limited branches remain in the result. Total model ceiling is one billion transitions; total active diagnostic ceiling is ten minutes, checked between bounded habitat chunks. An overrun is recorded as budget-ended rather than complete. Replay verification costs are separate.

Before collecting results, freeze this method, the audit source, runner, and tests alongside the hashes of all 47 prior scientific sources. Persist completed habitat chunks atomically. Publish a read-only view of the fixed examples and summary. Preserve earlier data and live saved journeys.
