# v0.12 code review

This is the first focused retrospective review of the current nursery and its shared evidence service. The prior project had automated tests and replay checks; this pass traced failure paths and checked whether implementation faults could be mistaken for research limitations.

## Findings and fixes

1. **P1 — canceled plans were missing from saved journeys.** In inquiry-app.mjs, choosing a manual intervention cleared the current plan immediately, while saves recorded only subsequent physical actions. Canceling that intervention before movement left no record of the plan change. On exposed seed 90101 at tick 28, resumed care moves west while an action-only replay tries to inspect; reload can reject and replace the journey. The new journey.mjs records control events independently of physics, replays trailing cancellations and preserves legacy manual-action semantics during migration. If replay fails, it attempts to retain the unverified original checkpoint in local storage before replacing the live journey. This affects interactive replay, not the frozen automated evaluation.

2. **P2 — exporting did not stop manual movement.** Export set the paused flag, but the animation loop advances whenever a manual journey exists. Export now ends that manual movement before saving the snapshot. A browser regression checks that the exported tick remains the live tick. Automatic pending care remains available to resume.

3. **P1 — one failed save disabled later saves in two older stores.** StudyStore and NurseryStore chained every write directly onto the previous promise. A transient write failure left that chain rejected, preventing subsequent attempts even after the filesystem recovered. They now recover the queue before attempting the next atomic write, while the failed call still reports its error. Regression tests obstruct the temporary path, confirm the previous durable file survives, remove the obstruction and verify a successful retry. LabStore already handled this correctly.

4. **P2 — a late final checkpoint could claim completion.** LabStore checked completed state before the active-time ceiling. A final result that arrived between timer checks could mark a run complete despite exceeding its limit. A reproduction with one millisecond remaining and a two-millisecond result returned complete. It now retains the measured work but marks the run budget-limited; invalid elapsed times are rejected before mutation. The newest result panels also require service completion before presenting a final conclusion. Existing audited v0.10/v0.11 runs were within their ceilings and are unaffected.

## Scientific and infrastructure review

Reviewed the active two-bed physical step, forecast step, observation interface, calibration, candidate model updates, proposal validation, ordinary and timely planning, reading-withholding comparisons, diagnostic branching, study summaries, replay, budgets and persistence. Checked the server's loopback binding, origin/host checks, explicit public-file surface, authenticated worker path, interruption accounting and worker dispatch. This was not a penetration test.

The forecast integrator intentionally assumes legal candidate actions and constant weather during each forecast; real blocked actions are passed to the estimator as waits. New tests exercise legal travel and every mechanism for 480 ticks in each of five conditions, crossing both fault windows and retaining water in transit. No physics/forecast mismatch was found in those checks. This does not validate forecast accuracy when learned parameters, unobserved state or future weather differ.

The new 48-tick fitted diagnostic is checked against frozen ordinary care for identical actions, care outcome and forecast cost. Tests verify that the coefficient reference does not read hidden moisture or future schedules, while the full-state reference is separately labeled. All forty prior scientific sources and completed report hashes remain unchanged; fixes are in the interactive application and service layers. Historical reports are evidence of the code that actually ran and are not silently rewritten.

## Limits and follow-up

The nursery's large UI module and several copied planner implementations make future divergence easier. This pass added a small shared journey/replay module where a demonstrated fault required it. A broad rewrite of frozen research implementations would weaken reproducibility; extract shared implementations only in a newly versioned experiment with parity tests.

This was not an exhaustive line-by-line review of every older maze, image, neural-training adapter or vendored dependency. Their existing regression suites and source hashes were checked; deeper algorithm reviews remain separate work. Store writes are atomic within each file, not a transaction spanning all three stores. The application remains a local research tool.

See verification (full source archive), [diagnostic method](v12-method.md), and [findings](v12-findings.md).
