# v0.12 findings: what limits the next decision?

The completed diagnostic gives a reason to audit model predictions before increasing planning depth. Supplying known current physical coefficients improved mean care by 1.46 percentage points at the existing horizon, while doubling the forecast horizon changed fitted-model care by −0.43 points and nearly doubled branch computation. Both descriptive intervals include zero. The result does not establish model error as the main limitation, and ensemble care remains the default.

This is Stage 1 of the [roadmap](../ROADMAP.md). We have completed the planned model/planning comparison and a focused retrospective code review. Reliable learned improvement remains unproved.

## Same checkpoint, five continuations

The [frozen method](v12-method.md) uses twelve already exposed v0.11 base seeds, five conditions and two clock-selected checkpoints per condition. Each of the 120 checkpoints supplies the same physical state and memory to five 96-tick continuations. All branches use the ordinary action menu, loss and tail controller, with the same two-million model-transition ceiling.

The known-dynamics reference receives current delivery, evaporation, delay and refill coefficients before each physical step. It keeps ordinary local observations and estimated hidden state. It does not receive future weather or change schedules. The full-current-state reference additionally receives current moisture, reservoir and pipe contents. These are offline observer grants, not capabilities learned by live Lumen or proven optimal controllers.

| Condition | Both beds viable | Branch model transitions | Budget-limited branches |
| --- | ---: | ---: | ---: |
| Fitted models, 48-tick horizon | 83.64% | 69,174,183 | 0/120 |
| Fitted models, 96-tick horizon | 83.21% | 137,441,321 | 0/120 |
| Known dynamics, 48-tick horizon | 85.10% | 66,924,010 | 0/120 |
| Known dynamics, 96-tick horizon | 84.57% | 131,838,931 | 1/120 |
| Full current state, 48-tick horizon | 85.16% | 66,411,054 | 0/120 |

All branches, including the limited branch, contribute their complete 96-tick outcome. When another plan cannot be reserved, the branch waits within its remaining physical window. Equal ceilings do not imply equal actual computation.

| Paired change | Mean care difference | 95% descriptive interval |
| --- | ---: | ---: |
| Known dynamics at 48 ticks | +1.46 pp | −0.72 to +3.63 pp |
| Known dynamics at 96 ticks | +1.36 pp | −0.44 to +3.16 pp |
| Longer horizon with fitted models | −0.43 pp | −1.41 to +0.55 pp |
| Longer horizon with known dynamics | −0.52 pp | −1.59 to +0.55 pp |
| Model × horizon interaction | −0.10 pp | −1.41 to +1.22 pp |
| Full current state beyond known dynamics | +0.07 pp | −0.36 to +0.50 pp |

Differences are averaged within base seed, then across twelve seed clusters. The intervals describe variation in this exposed panel; they are not confirmatory population-generalization intervals. No fresh validation cases were drawn. Raw percentages are not directly comparable with v0.11's full-season outcomes.

The coefficient benefit is uneven: seed-level differences at 48 ticks range from −2.92 to +9.58 pp. The small full-state difference does not establish that sensing is unimportant. Likewise, the longer-horizon result concerns this forecast, loss and tail heuristic; it does not rule out better planning methods.

## Inspectable behavior and code review

**Inspect the bottleneck** beside the nursery controls opens synchronized recorded habitats. Compare fitted 48- and 96-tick planning with a selected observer reference, scrub the same clock, jump to the first action difference and export all five continuations. The two examples are the first checkpoints in traversal order, not selected winners. Playback reads recorded frames and adds no simulation actions.

The [code review](v12-review.md) found and fixed four reproducible faults: canceled plans missing from saved journeys, manual movement continuing during export, failed save promises preventing future writes, and late final results being marked complete beyond their time ceiling. The review concentrated on the current nursery and shared evidence service. It was not an exhaustive review of every historical maze, image or neural-training implementation.

102 tests and both HTTP suites passed. Legal physical and forecast steps agreed across 2,400 checked transitions covering all five conditions and both fault windows. The 48-tick fitted branch reproduced frozen ordinary care. Browser checks covered replay, export pausing, reload continuity, reference selection and desktop/phone layouts. Independent replays reproduced all five exported diagnostic branches, the legacy ordinary controller's actions and cost, and a 230-action browser journey containing a cancellation event. See verification (full source archive) and replay evidence (full source archive).

## Cost, preservation and next decision

The diagnostic used **90,504 measured physical actions**, zero uncertain actions, **625,364,899 model transitions** and **393.1 seconds of active worker time**. Physical cost comprises 4,104 calibration actions, 28,800 collection actions and 57,600 continuation actions. Active worker time excludes browser rendering and service overhead. The engineering pilot, tests and independent replays are separate verification work, not additional validation cases.

The audit (full source archive) verifies all 42 frozen scientific sources, the method and pilot, the complete report and earlier evidence. All forty prior scientific sources remain unchanged. The full report is bottleneck-study-678090d1-bced-4849-8ad9-dc4f365e2885.json (full source archive); its provenance is in the manifest (full source archive).

Next, audit calibration and multistep prediction accuracy on common action sequences. Separate parameter error from stale estimates and changing weather before choosing one targeted revision. The current evidence does not justify spending nearly twice as many model transitions on the longer horizon. Any revised learner must then earn its place through a newly frozen, fresh full-care comparison under declared budgets. Better predictions alone will not satisfy the Stage 1 care-improvement gate.
