# v0.18: less repeated work, and a more specific replanning failure

Sharing the two beds' waiting projections preserved every tested score, action, trajectory and computation stop. Median paired search time fell **17.18% in steady and 21.27% in compound**, with **25.87% and 27.85% fewer executed model transitions**. This passes the frozen local speed criterion. The optimization is retained offline; it does not change the live controller or rescue the previously rejected two-decision policy by itself.

The first replanning forks supply a separate finding: withholding subsequent moisture and tank readings changes **none of the four proposed journeys**. Committing the old forecast through its next mechanism operation repairs the retained failure, but harms two successful controls. We have a local inconsistency to explain, not a general case for ignoring readings or following old scripts.

## Exact reuse without a persistent cache

Both beds' tail rules begin from the same fitted state and predict repeated waiting. They ask for moisture at different travel-plus-lag times. The shared implementation advances one cloned state to the longer endpoint and retains each bed's moisture at its requested tick. All models still have separate forecasts. No key rounding, approximate match or reuse across different starting states occurs.

The temporary memo holds two moisture values and one working model state, and expires before the next tail decision. Requests longer than 64 ticks use the original independent projections. Tests cover that fallback and unequal/equal endpoints. No exposed benchmark required the fallback.

Logical requests remain exactly equal to the frozen planner's charges. The existing controller budget uses these original charges, including reused work, while the report separately counts actual model executions. Consequently the optimization cannot silently give the candidate more search before a budget stop.

## The fixed timing comparison

Each retained checkpoint received two warmup pairs and ten timed pairs of complete twelve-root searches, alternating execution order. Explicit GC occurred before each search, outside its timer. Every root result—including forecast actions, states, groups, scores and logical steps—matched exactly. Warmups and both implementations' executions are charged.

| Fixed search | Original executions | Shared executions | Reduction | Original median time | Shared median time | Median paired time reduction |
|---|---:|---:|---:|---:|---:|---:|
| Steady, tick 113 | 70,215 | 52,051 | 25.87% | 67.23 ms | 56.18 ms | 17.18% |
| Compound, tick 115 | 115,199 | 83,116 | 27.85% | 86.32 ms | 68.03 ms | 21.27% |

The last column is one minus the median of paired time ratios; it is not calculated by dividing the two independently summarized median times. The declared criterion required at least 5% lower median paired time in both cases, reduced executions, and exact search/trajectory parity. It passed.

These are two exposed searches on this Mac, not a general application benchmark or energy measurement. Heap samples are available in the evidence view. They measure allocated heap immediately before and after a search, including returned results and runtime behavior, and can change with garbage collection. They do not establish peak-memory savings. The structural memo bound is more informative than claiming a memory reduction from those samples.

All four earlier 48-tick candidate continuations and both retained complete candidate seasons reproduced their frozen actions, frames, logical work and outcomes exactly. Retained branching plans also matched. The compound season still begins its original budget waiting at state tick **437**. This confirms semantic and allowance parity; no changed-budget policy was evaluated.

## What changes at the first fork

The four cases use the two retained v0.17 checkpoints, each with its original journey and the previously lowest-realized-cost alternative. Alternatives are chosen with hindsight from earlier evidence and are not available to the learner. Current results did not determine case selection.

A shadow memory shares the initial observation, physical actions, equipment and weather, but receives no later moisture or tank readings. It advances alongside the original physical trajectory only until the first divergence. We inspect the responsible planning decision, propose one journey from each memory, and execute matched interventions from the same physical checkpoint. All physical branches use real local readings and the actual memory. The shadow supplies only a proposed journey.

| Case / initial journey | Decision tick | Current choice | Choice without later readings |
|---|---:|---|---|
| Steady / A valve → 50% | 115 | Observe here | Observe here |
| Steady / B valve → 50% | 122 | Observe here | Observe here |
| Compound / Read bed A | 119 | Read bed A | Read bed A |
| Compound / B valve → 0% | 116 | B valve → 50% | B valve → 50% |

At the failed steady case's tick 115, the ordinary planner scores observing here at **5.16724**, just below opening B to 50% at **5.17101**. Yet both its original forecast script and the existing heuristic applied to its current selected model propose traveling to B and opening it. A later moisture reading is therefore not necessary to produce this particular disagreement between the heuristic follow-up and the actual search rule. Weather, belief estimates and other state still differ over time; this does not establish that all model updates or readings are irrelevant.

## One repaired failure, two damaged controls

Each intervention commits one proposed journey, or the script through its next valve/shade/drain operation, capped at sixteen ticks. Ordinary care resumes afterward. The current-choice commitment arm checks whether forcing a prefix alone changes the result; it exactly matches ordinary outcomes in all four cases.

| Case | Common suffix | Ordinary | Commit current choice | Choice without later readings | Current model's next operation | Original script's next operation |
|---|---:|---:|---:|---:|---:|---:|
| Steady, original journey | 94 ticks | 69.15% | 69.15% | 69.15% | **100.00%** | **100.00%** |
| Steady, successful alternative | 87 ticks | 100.00% | 100.00% | 100.00% | 100.00% | **97.70%** |
| Compound, original journey | 92 ticks | 100.00% | 100.00% | 100.00% | 100.00% | 100.00% |
| Compound, successful alternative | 95 ticks | 100.00% | 100.00% | 100.00% | **98.95%** | **98.95%** |

Percentages measure both-bed viability over the same remaining suffix within each case. They differ from earlier 96-tick metrics because the common prefix is excluded. No combined average or new confidence interval is assigned to this deliberately selected diagnostic set.

In the failure, committing the five-action journey to open B changes realized suffix cost from **119.1204 to 0.2479**. This intervention bundles travel, time, readings and the valve operation; it does not isolate the valve's physical effect alone. In the successful steady alternative, the old script closes B after one waiting tick while the current heuristic waits four ticks before the same journey. Using that stale timing causes a small viability loss. In the successful compound alternative, moving to A as either tail proposes is worse than immediately reopening B as ordinary planning chooses.

The distinction matters: a useful forecast can be abandoned without a new reading causing that choice, while abandoning a forecast can also protect a successful trajectory. A blanket commitment rule fails these controls.

## Costs and verification

| Work | Physical actions | Logical model requests | Executed model transitions | Active time |
|---|---:|---:|---:|---:|
| Paired search benchmark, including warmups | 0 | 4,449,936 | 3,846,972 | 3.66 s |
| Four short and two full trajectory checks, including calibration | 1,494 | 24,611,312 | 17,292,105 | 15.43 s |
| Four replanning cases and all five interventions each | 2,224 | 13,649,456 | 13,649,456 | 7.32 s |
| Development total | **3,718** | **42,710,704** | **34,788,533** | **26.41 s** |

The fork diagnostic uses the unchanged ensemble, so its requested and executed transitions are equal. All reservations closed. All earlier 96-tick baselines and their ordinary suffixes reproduced exactly. No fork branch reached its model limit.

The audit (full source archive) verifies all 61 scientific source hashes, method-before-run ordering, prior evidence, panel coverage and accounting. Independent replay of the retained original-choice failure and compound control reproduced the complete fork records exactly, using **1,122 physical actions and 7,590,371 model transitions** over 4.08 active seconds. This verification is separate from the development totals. All **133 automated tests** and both HTTP integration suites passed. The delivery report (full source archive) records exact durable exports, all sixteen case/intervention combinations, immediate playback and end restart, expanded evidence, desktop/mobile layout, and preserved nursery navigation. The live season remains paused at tick 230/480.

**Where the plan changes** beside the nursery controls opens `/replanning.html`. It exposes timing pairs, logical versus executed work, local readings, fitted models, competing scores, all five outcomes, synchronized intervention replays and exact exports. The live engine hash and ensemble default remain unchanged.

## Next gate

The exact speed improvement creates a specific opportunity: a separately frozen complete-season development comparison that charges executed model transitions against the same fixed physical and computation ceilings. Keep the two-decision search rule, models, loss and parameters unchanged; change only the implementation and explicit accounting. Count fitting and every rejected branch. Compare all conditions against ensemble care and require the full-season gate before any fresh evaluation or live use. The previous pre-budget care failures remain a reason the candidate may still lose.

Keep these fork cases as development controls. If additional useful search does not repair full-season behavior, focus the next design on consistency between the forecast continuation and the executed decision rule, without rewarding operations that are repeatedly deferred. Do not replace that diagnosis with a blanket instruction to commit or ignore readings.

Stage 1 remains active. This cycle demonstrates a measured computation improvement and narrows a local decision failure; it does not yet demonstrate reliable improvement from accumulated experience.
