# v0.19: spend only executed model work

Freeze this method, the two new scientific modules, tests and runner before new development outcomes. Preserve all 61 earlier scientific sources and their evidence. Work uses only this Mac and the three already exposed pilot seeds. Nothing here changes the live nursery.

## Fixed question and comparison

Does charging actual executed model transitions let the unchanged shared two-decision planner improve complete-season care within the existing allowance? Use seeds 93101, 93103 and 93113, the existing scenario offset 9,100,001, and all five conditions: steady, pipe, dry, supply and compound. Each seed receives one shared 342-action measured calibration. Run 480-tick seasons in three arms, in fixed order:

1. Legacy allowance: v0.18 shared planner, still charged for logical requests.
2. Executed allowance: the identical shared planner charged for actual model transitions.
3. Ensemble care: the frozen live comparator with its original accounting (all its transitions execute).

Keep the models, fitted inputs, loss, reserve valuation, horizon, observation groups, root menu, journey completion/replanning and tie ordering fixed. Change only which work count the allowance reads. Retain the 12,000,000 per-season ceiling, 2,000,000 per-search reservation and 100,000 update reservation. The conservative rule stops new searches when current charge plus both reservations exceeds the ceiling. Remaining ticks still observe, update models and wait. Do not reclaim this reserved headroom, deepen search, commit old tails or retune parameters.

## Work ledger

The frozen planner continues accumulating logical requests. Actual transitions equal logical requests minus the cumulative shared-projection reuse count, subtracted once. Validate nonnegative safe integers, requests = executions + reuse, and that instrumented planning requests do not exceed total logical requests. All rejected root and child evaluations remain charged. Record searches, evaluated/unchosen roots, total search transitions, reused requests, projection groups, fallbacks and the existing two-scalar memo bound. The search total includes selected and rejected branches; it is not a per-branch cost attribution.

Separately count model-bank, fitter and proposal transitions outside search using the exact state sizes advanced by frozen rememberDiagnosis. Assimilate the current observation first (the subsequent wrapper call is idempotent), then count both banks plus their scratch models and any pair of proposal states. Require this independent count to match the shared planner's total-minus-search ledger. Ensemble has no reuse; its search count is total minus these outside-search updates. Regression arithmetic, allocation, copying and memo bookkeeping are not model transitions. Active wall time includes them; this experiment is not a new wall-time speed comparison or energy measurement.

Before freezing, seed 90101 tests exact pre-limit behavior and memory, accumulated credit beyond the old logical ceiling, actual reservation stopping, post-stop updates, invalid ledgers and completed-season no-ops. Tests are software verification, not development evidence.

## Bounded run, controls and decision

Run all 45 complete seasons once: 21,600 physical care actions plus 1,026 calibration actions, exactly 22,626 physical actions. Ceiling: 600 million executed model transitions and 600 active seconds. Reserve each calibration or season before execution and persist after it. An interruption leaves its reservation visible; a caught failure conservatively charges the outstanding physical reservation as uncertain. Any incomplete/failed/over-budget run fails the development gate. Do not rerun to reverse an outcome.

Require all 15 legacy and 15 ensemble rows to reproduce v0.16 metrics and logical charges exactly. Require each executed-allowance trajectory, actions and ledger to match legacy up to the first legacy-limited action (the whole season if never limited). Record the first different action, each budget stop, and common before/after-stop viability windows for all 15 pairs. These windows describe consequences of the accounting intervention; they do not establish that all pre-stop planning errors have been repaired.

The primary development gate is unchanged from v0.16: complete all cases and obtain nonnegative mean paired both-bed viability versus ensemble over the 15 matched habitats. Report per-seed and per-condition means, all rows, cost totals, and limited-season counts. Three reused seeds are development evidence, not a fresh generalization claim; do not use an interval over 15 correlated habitats to imply otherwise. A pass earns a separately frozen fresh evaluation, not promotion. A failure retires this accounting change as a sufficient care repair, while retaining exact reuse as a valid computation optimization. No after-the-fact changes to this rule.

## Delivery

Retain all three complete trajectories and cumulative work curves for the first seed's steady and compound cases, chosen before outcomes. Add a synchronized read-only replay with budget-stop and first-difference jumps, clear actual/logical accounting, full development comparisons and exact exports. Playback executes no model or physical transitions. Independently reproduce the executed and ensemble retained episodes using one new shared calibration for verification, charging that work separately (2,262 physical actions; at most 60 million model transitions and 120 active seconds). Audit all sources, prior evidence, result sums, controls, ceilings and freeze ordering; run appropriate tests and HTTP/browser checks. Update the roadmap based on the frozen gate.
