# v0.14 findings: journey ranking and the cost of dropping the ensemble

The fresh comparison **rejects selected-model scoring as an improvement**. It reduced complete-season both-bed viability by **2.25 percentage points** versus ensemble care; the paired 95% interval is **−4.19 to −0.30 pp**. Ensemble care remains the live default. Stage 1 is still active.

## What the exposed audit found

The [fixed ranking method](v14-method.md) collected 60 plan-boundary checkpoints from six previously used seeds, across five conditions. All twelve ordinary journeys were tested twice from identical states: execute the selected model's full forecast script, or commit the initial journey and then resume ordinary ensemble care. This produced 1,440 physical branches.

| Initial journey score | Lowest measured running cost | Mean cost regret | Both beds viable in the next 48 ticks |
| --- | ---: | ---: | ---: |
| Existing weighted model score | 25 / 60 | 4.2067 | 87.60% |
| Selected model alone | 22 / 60 | 1.6475 | 88.92% |
| Weighted models with recorded future weather | 15 / 60 | 4.3725 | 87.57% |

The selected-model score reduced the severity of mistakes without selecting the exact minimum more often. Its +1.32 pp viability difference was mostly attributable to one seed, 93101. It was worse on 93139 and unchanged in viability on the other four. These correlated, exposed checkpoint results motivated a candidate, not a claim of generalization.

The live planner imagines a heuristic follow-up routine, then uses full replanning during actual care. Every tested candidate's ordinary continuation eventually differed from the imagined action sequence. Replanning reduced running care cost by 11.35 units on average relative to physically executing that fixed script; mean absolute difference was 13.34. Even exact physical script outcomes selected the lowest-cost ordinary continuation at only 21/60 checkpoints, with mean regret 0.8548.

This is a real mismatch, but not proof that the heuristic tail is the dominant error. Selected-model running-cost prediction error averaged 22.75 against the physical script and 19.68 against ordinary continuation: interactions sometimes make the changed continuation closer to the prediction. Predicted scores also include a terminal reserve estimate, while observed cost covers only the next 48 ticks. No additive causal attribution is justified.

The first retained steady checkpoint makes the problem visible. At tick 113, the planner chose “A valve → 50%.” Its measured ordinary continuation had cost 83.53 and 56.25% both-bed viability. “B valve → 100%,” ranked last by the forecast, had measured cost 0.13 and 100% viability. This is a retrospective example within a fixed menu, not an available foresight policy. Other cases show negligible differences despite different action sequences.

## What survived the fresh test

The [frozen care method](v14-care-method.md) changed only journey scoring. The candidate retained the entire hypothesis bank, fitting, validation and switching, but used the active model's cost instead of mixing three models. Both arms began at tick zero and could interrupt or replan normally. Twelve fresh seeds × five conditions × two arms produced 120 complete seasons.

| Controller | Complete-season both-bed viability | Model transitions | Limited seasons |
| --- | ---: | ---: | ---: |
| Selected-model scoring | 78.94% | 51,197,326 | 0 / 60 |
| Ensemble scoring | 81.19% | 151,923,004 | 0 / 60 |

The candidate's average paired difference was −2.2465 pp, with 95% interval −4.1902 to −0.3029 pp. Nine of twelve seed averages were negative. The preregistered promotion rule required at least +1 pp and a positive lower interval; it failed. The roughly 66.3% reduction in forecast/update transitions does not compensate for worse care, and this study did not preregister a noninferiority claim. This transition count is not a measurement of energy savings.

Descriptively, the largest condition difference was under pipe restrictions: −8.54 pp. Steady conditions were +0.66 pp; dry −0.68 pp; supply −0.03 pp; compound −2.64 pp. These condition summaries are not separate confirmatory tests. Do not tune the candidate on these now-exposed evaluation cases and call them fresh again.

Equal-cost detail: the diagnostic retained the existing ensemble order when selected-model scores tied. The candidate preserves the ordinary menu order. This changes the tied selection at four development checkpoints, with no change in their aggregate viability. Neither tie rule was tuned to physical outcomes. All prefixes and score calculations are retained for inspection.

## Cost, preservation and verification

| Work | Physical actions | Model transitions | Active time |
| --- | ---: | ---: | ---: |
| Exposed ranking diagnostic, including calibration and collection | 85,572 | 261,583,873 | 149.35 s |
| Fresh complete-care comparison, including calibration | 61,704 | 203,135,954 | 135.44 s |
| Independent replay verification | 4,224 | 12,974,756 | Recorded separately from research |

No uncertain or budget-limited research actions were recorded. A freeze preflight initially encountered two files with the same basename; it stopped before any research computation, and its log and partial copy remain preserved. The runner then froze source paths correctly. Unit tests and isolated HTTP checks are separate verification work, not extra evaluation cases.

All **115 tests** and both HTTP suites passed. The latter verified durable checkpointing, stopping, conservative interrupted reservations, private-file protection, cross-origin rejection and worker authentication. Exact replay reproduced all 48 retained diagnostic branches, both instrumented forecast sets, and all four retained full seasons. All 51 scientific sources match their frozen hashes; earlier sources and evidence remain intact.

The audit record (full source archive), replay verification (full source archive) and delivery verification (full source archive) retain the checks and file hashes. The controller hash is `4780383408a8f85c47d6bd429e16a7a9dcfb8ca92ab999dc00b377f012ec442c`.

## What changed in the application

**Why this journey?** beside the nursery controls opens `/rankings.html`. It includes every rejected candidate's numerical outcome, the two fixed exposed examples, and three replay comparisons: prediction versus the same physical script, script versus ordinary replanning, and the planner's choice versus another tested journey. Play/pause, shared time, first-difference jumps and exports work for the diagnostic and fresh complete seasons. Mobile layouts stack the habitat cards and contain tables in horizontal scrolling regions.

The selected-model candidate remains a research implementation. The live saved journey and ensemble default are preserved. Existing nursery, maze and prediction pages remain linked.

## Next bounded question

**Why does subsequent replanning sometimes undo the value of a good initial journey?** Preserve ensemble uncertainty handling. At exposed failure checkpoints, record the predicted next operation and its urgency, then the next actual plan, local reading and model change. Locate the first harmful reversal or repeated deferral, distinguishing state error, changing model support, terminal reserve valuation and the assumed follow-up. Include successful controls and a finite action budget.

Use that trace to choose one bounded planning-consistency correction, potentially an explicit second decision in the forecast if the evidence supports it. Do not simply lengthen the horizon, remove the ensemble, or scale the run. A correction must still face fresh complete seasons. The v0.14 result demonstrates why short-window ranking gains are insufficient for promotion.
