# Conservatory v0.9 — completed findings

The clearest gain is cheaper image alignment: the coarse search preserved 60/60 successes while reducing predicted pixel-channel values by **95.18%**. The richer maze and nursery tests exposed limits: fresh gate learning was more reliable than transferred timing priors, and the learned investigation selector did not improve nursery care.

All four frozen suites completed on this Mac. The audit (full source archive) verifies **214,838 physical actions, zero uncertain actions**, the 35 frozen source files, and preserved earlier source and evidence. The photographic pilot is separate: another 258 actions and no human ratings. See the [method](v9-method.md) for protocols and limits.

## Nursery: asking fewer questions did not improve care

Six training habitats supplied **288 branch outcomes**. The frozen selector then faced twelve separate habitats, five conditions each. Compound disturbances were excluded from selector training.

| Policy | Both beds viable | Test starts / season | Test actions / season | Evaluation model transitions |
|---|---:|---:|---:|---:|
| Learned investigations | 79.79% | 0.68 | 6.37 | 151,778,855 |
| Ensemble care | 80.44% | 0 | 0 | 156,617,026 |
| Decision-value tests | 80.53% | 1.33 | 19.33 | 216,376,713 |
| Shuffled training outcomes | 80.47% | 0.02 | 0.13 | 156,699,935 |

The primary learned-minus-ensemble difference was **−0.65 percentage points**, with a paired 95% interval of **−1.64 to +0.35 pp** across twelve base seeds. No care advantage was established. Ensemble care remains the default; the learned selector is available as an optional experiment. The shuffled control almost never investigated, which also cautions against interpreting fewer tests as successful learning.

The 36 matched diagnostic checkpoints help distinguish mechanisms:

- A tested investigation improved on ordinary ensemble care at **8/36** checkpoints. The v0.8 initial choice missed six of these eight opportunities.
- For the retrospectively best tested continuation, withholding moisture/tank readings reduced viable time at **5/36** checkpoints. In several other cases, the physical intervention accounted for the improvement without an incremental benefit from those readings.
- The full-state planner improved on ensemble care at **7/36** checkpoints. It is a reference using the same bounded planner, not an optimal oracle.

These are retrospective diagnostic comparisons, not independent estimates of selector performance. A test can be informative without beating ordinary care, and can physically help without teaching a useful fact. The learned selector's training labels combine both effects, which remains an unresolved limitation.

Nursery accounting: 6,156 calibration actions + 46,080 training collection/branch actions + 115,200 evaluation actions + 34,560 diagnostic branch actions = **201,996**. All stages used **1,083,966,738 model transitions** and **98,784 neighbor comparisons**. The service recorded 645.9 seconds of active computation; this is a run measurement, not an energy estimate.

## Maze: familiar paint is an unreliable shortcut

All 48 generated evaluation rooms had a verified physical witness. Training used individual gates; evaluation introduced alternate passages, distant connections, changed paint, sparse obstacles and larger boards.

| Policy | Goals reached | Mean actions, including failures | Evaluation search transitions |
|---|---:|---:|---:|
| Repair transferred timing priors | 45/48 | 30.38 | 369,848 |
| Learn afresh | **48/48** | 31.58 | **258,672** |
| Replan every action | 45/48 | 27.65 | 698,092 |
| Keep timing assumptions | 45/48 | 29.88 | 397,732 |

| Condition | Repair priors | Fresh | Replan | Keep assumptions |
|---|---:|---:|---:|---:|
| Near connections, stable paint | 12/12 | 12/12 | 12/12 | 12/12 |
| Near connections, shuffled paint | 12/12 | 12/12 | 12/12 | 11/12 |
| Distant connections, stable paint | 12/12 | 12/12 | 11/12 | 12/12 |
| Distant connections, shuffled paint | **9/12** | **12/12** | 10/12 | 10/12 |

All three repair-policy failures occurred with distant connections and shuffled paint, and were search-limited. A lower mean action count is not a win when episodes fail early. The uncached policy can choose different routes, so its results do not isolate caching alone. These twelve-seed results apply to this room family; they do not establish general navigation ability.

The separate novelty curriculum retained **6/18** generated children. It skipped donor trials for ten cheaply solved residents, made two actual donor attempts, and recorded **zero winning transfers**. The remaining non-skipped cases had no distinct eligible palette donor. The novelty filter now rejects redundant behaviors, but this run supplies no evidence that transfer is helping on the harder family. Accepted and rejected children remain inspectable, and the interface can load the recorded learner's starting palette with its room.

The comparison charged **8,004** physical actions; the curriculum charged **3,210**, including training, controller panels, resident/donor attempts and executed author witnesses.

## Images: the same numerical result with much less alignment work

| Strategy | Both criteria met | Mean actions | Predicted pixel-channel values |
|---|---:|---:|---:|
| Coarse → fine | **60/60** | **8.60** | **11,937,792** |
| Exhaustive search | **60/60** | **8.60** | 247,627,776 |
| Fixed control | 48/60 | 9.93 | 6,389,760 |

The coarse strategy used **95.18% fewer predicted channel values** than exhaustive search. It skipped alignment when the measured control already predicted an acceptable edit and searched at finer resolution only when necessary. Both movable strategies averaged target RMSE 0.00317 and outside RMSE 0.000366. Every final result, including fixed-control failures, satisfied the outside-preservation guard; fixed control failed the shifted targets.

This count measures evaluated pixel-channel values, not total CPU instructions, wall-clock speedup or energy. The gain is established on these five authored control conditions and their frozen image seeds. The full comparison cost 1,628 physical commands.

## Photographic review: ready to collect judgments

The separate numerical pilot passed **30/30** trials: three correlated crops from one NASA satellite montage × five conditions × two search strategies. It cost 258 actions and 259,565,568 predicted channel values. This small source set does not establish photographic or perceptual generalization.

The browser now supports a randomly ordered A/B review of imported photographs, separate intended-change and outside-preservation judgments, “Equal” and “Cannot judge,” source provenance, persistent review records and exports. Human ratings are **not yet collected**. The NASA example is attributed in the interface and third-party notes (full source archive).

## What is ready to use

- **Nursery:** Compare futures pauses the habitat and replays matched continuations with a shared time slider and an exportable replay checkpoint. New seasons can use ensemble care, decision-value tests or the frozen learned selector. The model used by a journey is saved with it.
- **Gate maze:** alternate passages, distant wiring, changed appearance, recovery from contradicted timing assumptions, bounded curriculum results, and replayable generated cases. Play starts without changing the room.
- **Image trial:** coarse/exhaustive search controls, local photographic input, drawn selections, numerical evidence and a separate A/B review.
- **Exports:** write a durable JSON or PNG file under `data/exports/` on this Mac, show its full location, and provide a downloadable copy. Browser navigation does not erase service studies or IndexedDB photo reviews.

The next useful work is to improve the *quality and allocation* of information: train nursery choices from a label that separates reading value from physical intervention, and diagnose why transferred maze priors spend their search budget before successful recovery. A broader photographic collection and actual independent judgments are needed before treating numerical image success as perceptual quality. More runtime alone does not answer these questions.
