# v0.13 findings: prediction improved; care advantage remains unproved

The calibration audit ruled out the first suspected problem. Weather was a substantial source of long-range prediction error, and a forecaster using observed weather history reduced that error in development. The fresh full-care comparison did not establish a care advantage. Ensemble care remains the default.

## What the audit found

On twelve exposed base seeds, the existing calibration recovered all 24 delays correctly. Its maximum relative gain error was approximately 0.000025%, and maximum relative evaporation error approximately 0.0000065%. Although calibration uses a short adaptive history, these cases did not support rewriting it.

The common-action audit used 60 complete ensemble-care seasons and 2,220 clock-selected checkpoints. Each forecast received the same recorded next 48 actions. This measures conditional prediction accuracy; those future actions are not known to the live planner. Predictions were produced before scoring them against the recorded outcomes.

| Forecast information | Mean moisture error at 12 ticks | At 48 ticks |
| --- | ---: | ---: |
| Weighted model mixture | 1.88 points | 7.28 points |
| Selected model | 1.16 points | 6.79 points |
| Exact current state | 0.81 points | 6.34 points |
| Exact state and current coefficients | 0.49 points | 5.37 points |
| Also reveal recorded future weather | 0.04 points | 1.05 points |
| Also reveal future coefficient changes | 0.00 points | 0.00 points |

The observer references progressively change what information is available. Their effects depend on this order and are not independent causal contributions. A weighted moisture prediction also differs from the planner's weighted nonlinear care loss. The full dynamic reference reproduced the physical trajectory exactly, supporting integrator correspondence on these legal action sequences.

This pointed to weather forecasting as one concrete change worth testing. A fixed pilot on the first three exposed seeds reduced 48-tick mixture prediction error from 6.97 to 6.01 moisture points, about 13.8%. At twelve ticks it changed from 1.76 to 1.64 points. No settings were tuned after that pilot.

## Fresh complete-care comparison

The [method](v13-method.md) and all 47 scientific sources were frozen before twelve new base seeds were run under steady, pipe, dry, supply and compound conditions. The weather candidate and ensemble baseline each completed 480 physical ticks per condition. Calibration, candidate models, local observations, action menu, loss, planning horizon and computation ceiling were shared.

The candidate fits a temporal recurrence separately to heat and rain. It checks each fit against recent held-back observations and falls back to current weather when the fit is unstable, uninformative or fails that check. It receives no simulator phase, future weather or change schedule.

| Controller | Both beds viable | Model transitions | Limited seasons |
| --- | ---: | ---: | ---: |
| Weather-informed care | 85.30% | 157,389,591 | 0/60 |
| Ensemble care | 85.00% | 160,715,887 | 0/60 |

The primary paired difference was **+0.30 percentage points**, with a **95% seed-cluster interval of −1.57 to +2.17 points**. It failed the frozen promotion rule: at least one percentage point mean gain and an interval wholly above zero. This is neither evidence of equivalence nor proof that weather forecasting cannot help.

The candidate used a temporal forecast for at least one weather channel in 7,679 of 8,897 plans. Its complete-care model-transition count was about 2.1% lower, but different actions and numbers of plans affect that total. This was not a preregistered noninferiority or efficiency claim. Per-seed care differences ranged from −7.33 to +5.42 points; the average does not describe every habitat.

## What is usable now

**Check the predictions** beside the nursery controls opens `/predictions.html`. Select the recorded case, bed and observer reference; scrub a moisture forecast against its actual continuation; compare errors at five horizons. Below, play the two fresh care journeys together, jump to their first different action and export the recorded pair. The first fresh seed's steady and compound cases are retained regardless of outcome.

The page links back to the nursery and maze. Browser checks confirmed that returning to the nursery preserves the existing paused journey at tick 230. The prediction plot stays readable on a narrow screen; care cards stack and tables scroll within the page. The temporal candidate is available in the research implementation and recorded replays; the live default remains ensemble care.

## Cost and verification

The fresh comparison used **61,704 physical actions**, zero uncertain actions and **318,121,102 total model transitions**, including shared calibration and **1,220,050 weather recurrence transitions**. It also recorded **2,411,888 regression rows**. Active worker time was **304.2 seconds**. These counts do not assume that one weather recurrence step costs the same CPU time as a two-bed forecast.

Separate exposed development work comprised:

| Work | Physical actions | Physical-model transitions | Weather recurrence steps |
| --- | ---: | ---: | ---: |
| Initial calibration audit | 4,104 | 15,624 | 0 |
| Common-action prediction audit | 32,904 | 154,321,320 | 0 |
| Fixed weather pilot | 8,226 | 36,353,242 | 40,115 |

108 automated tests and both isolated HTTP suites passed. Independent replay reproduced all four retained fresh seasons, their action sequences, care metrics and model costs. That adds 1,920 replay actions as verification work, not fresh validation. All 42 prior scientific source files and earlier completed reports remain unchanged.

The audit (full source archive) indexes the complete care report and verifies source, protocol, development records, action accounting and prior evidence. See also the manifest (full source archive), replay verification (full source archive) and software verification (full source archive).

## Decision and next question

Keep ensemble care. Do not tune the weather forecaster on these newly exposed evaluation cases or treat its development prediction gain as an achieved care gain.

Next, audit **action rankings**: when the planner favors one care journey, does that journey outperform the alternatives under matched continuations? Record forecast costs before outcomes, compare the same ordinary action prefixes under the real continuation policy, and distinguish ranking errors from differences introduced by the heuristic forecast tail. This tests the connection between prediction and decisions directly. Choose one bounded planning change only after that diagnosis, then require another frozen fresh comparison. Stage 1 remains active in the [roadmap](../ROADMAP.md).
