# Conservatory v0.9 — testing whether questions change outcomes

The nursery now branches an exact checkpoint into matched continuations. A small investigation selector learns from those outcomes. The maze separates gate appearance from timing and connection structure. Image steering uses a bounded coarse-to-fine alignment search and a separate photographic review. All run locally; these are authored environments and bounded controllers, not evidence of AGI or sentience.

## Frozen evaluation

`experiments/frozen/conservatory-v9/manifest.json` records all 35 scientific sources, their combined hash, protocols, seeds and budgets before evaluation. The v0.8/v0.7/v0.6/v0.5 scientific sources and earlier evidence remain unchanged. UI files may change without changing a scientific run. The audit checks both versions and action accounting.

### Nursery: matched branches

A snapshot includes the physical state, recent observations, model bank and the simulator's future weather schedule. Pending plans are cleared identically in the branch copies. The live habitat is paused and unchanged. Controllers receive ordinary local observations; only the simulator and explicitly labeled observer reference use hidden state.

Each branch runs at most 96 physical ticks. Ordinary ensemble care is compared with four bounded tests: delivery at A, delivery at B, shared supply, and isolation followed by a supply reading. A test first executes its fixed travel/intervention/inspection prefix (at most 24 ticks), then resumes ensemble care. Another branch repeats each prefix while withholding moisture and tank readings through the first following care decision. Equipment observations remain available. This estimates the incremental value of those moisture/tank readings, conditional on the same physical test; it does not measure every possible kind of information.

A full-state planner receives the true current physical state and coefficients. It is a privileged reference using the existing planner, not an optimal oracle or a proven upper bound. The v0.8-choice branch compares its initial macro choice followed by ensemble care, not the entire v0.8 policy. “Best tested” is selected retrospectively among these five continuations. The latest interactive comparison, including its complete replay checkpoint, persists in IndexedDB and can be exported to a durable Mac file. Interactive comparisons charge up to 960 actions; interrupted computations charge their reserve as uncertain.

### Nursery: learned investigation selection

Six training base seeds supply ordinary-care moments at or after ticks 112, 224 and 336, when a pending journey ends. This clock sampling never uses a hidden disturbance label. Training excludes compound faults. Each candidate is labeled by its 96-tick viable-time difference against ordinary ensemble care. Physical intervention effects and the cost of postponing care are included in this label. Diagnostic reading-withholding branches are separate from training labels.

The selector is a small nearest-neighbor regressor: nine same-test examples, standardized observed/belief features, inverse-distance weighting, a 0.5 percentage-point minimum predicted gain and an unfamiliarity cutoff of 2.5. Features include moisture estimates, reading ages, valve/shade beliefs, tank estimate, disagreement, mismatch indicators, travel length, remaining time and current heat. It can propose an investigation only after a suspected mismatch, with a 48-tick cooldown and at least 96 ticks remaining. Tests may be interrupted by model revision; reported test counts are attempted starts and test-action counts are actual ticks. It does not acquire a new action vocabulary.

After training, the selector stays fixed across twelve held-out base seeds × five conditions × four policies: learned, ensemble, v0.8 decision-value, and shuffled-label learned selection. The shuffled control rotates gains within each test category, preserving marginal gains but disturbing their feature association. It is secondary and exploratory. Calibration is performed separately for each habitat and charged: 182 measurement actions plus 160 normal-validation actions.

The preregistered primary contrast is learned minus ensemble viable time, averaged across five conditions within each of twelve base seeds. Its paired 95% t interval uses 11 degrees of freedom. Individual conditions are correlated, not 60 independent replicates. Training, calibration, policy evaluation, branch diagnostics, model transitions and neighbor comparisons are all reported. Maximum physical budget: 201,996; maximum active time: 30 minutes. A learned policy choosing few or no tests is a possible result, not proof of successful learning.

### Maze: appearance, wiring and bounded planning

Generated boards vary between 15/17/19 columns and 11/13 rows, with reflection, sparse obstacles, goal placement and two or three alternative gated passages. Plates and nearby crates share the starting side, and near or cyclic distant wiring chooses which passage they control. Stable versus shuffled paint changes only appearance, using an independent random draw; physics and geometry remain paired. Training uses one gate at a time. This is a broader but still deliberately restricted family of rooms.

A palette learned in training supplies timing priors, but each gate has a separate working rule. Observed activation identifies connections. Release tests measure timing. Contradicted priors are discarded; when a crate prevents release, the learner physically removes the support before measuring. Four policies compare repair, learning afresh, uncached replanning, and keeping assumptions. All receive 180 physical actions and 30,000 search transitions. The uncached policy can choose different routes; this comparison is not an isolated measurement of caching overhead. Budget-limited failures remain in the denominator, and mean actions include failures.

Twelve independent base seeds × four paired appearance/wiring conditions × four policies produce 192 trials. An independent author witness uses hidden connections to execute a physical solution where possible. Its moves and search transitions are charged separately and included in total physical actions; the witness never enters learner memory. Maximum physical budget: 49,680.

The separate curriculum runs three seeds × six generations. It mutates passage count, wiring, paint, obstacle density and room seed. A fixed panel (fresh, founder-repair, founder-assumption) creates a six-value behavioral profile from success and normalized cost. Admission requires a physical witness, a successful learner in 12–150 actions and Euclidean profile novelty of at least 0.08. A resident solving within 28 actions skips donor trials. Otherwise up to two distinct palette donors are ranked by profile difference and tried. Every panel, resident, donor, rejected child and witness is charged. Maximum: 24,300 physical actions. This is a bounded adaptation of generation/transfer/novelty ideas, not a reproduction of Enhanced POET, and selected children are exposed training data.

### Images: computation and photographic review

The coarse search first solves with the measured control at its current position. If its predicted target and outside errors already satisfy the guard, no alignment search runs. Otherwise, 45 coarse positions use every fourth pixel in each dimension, followed by two 3×3 full-resolution refinements (two-pixel then one-pixel steps). A moved control is physically re-probed; both probe edits and their undos are charged. The old exhaustive strategy is preserved inside the new engine as a reference and checked against v0.8.

Twelve frozen seeds × five control conditions × three strategies (coarse, exhaustive, fixed) yield 180 procedural trials. Both movable strategies use the same initial calibration, proposal solver, preservation guard and maximum 64 commands/eight proposals. Search grids differ; this is a comparison of complete search strategies, not equal candidate budgets. Counts are predicted pixel-channel values, including the search, not FLOPs, memory traffic, latency or energy. Numerical target/outside criteria are unchanged. Maximum physical budget: 11,520.

The separate photographic pilot uses three correlated 128×128 crops of one [NASA Earth satellite montage](https://science.nasa.gov/photojournal/earth/), credited to NASA under its [media guidance](https://www.nasa.gov/nasa-brand-center/images-and-media/). Five conditions and two search strategies give 30 numerical trials. These are not thirty independent photographs. Pixels, crop coordinates, color conversion, mask, source hash, all actions and outcomes are saved. The results establish only this small numerical check.

The browser can import other photographs locally, draw a selection, run a randomly ordered A/B comparison, and record intended-change and outside-preservation judgments separately. Controller identities and numerical results appear only after rating. “Equal” and “Cannot judge” are explicit choices. IndexedDB retains each review and its provenance; export also includes unrated evidence. No human judgments are fabricated. This is a review tool and pilot, not a validated perceptual metric or multi-rater study.

## Research connection

[CausalWorld](https://arxiv.org/abs/2010.04296) motivates separating intervention dimensions and checking generalization across controlled changes. Here those dimensions are paint, wiring and geometry; we do not implement its robotic manipulation benchmark. [Enhanced POET](https://arxiv.org/abs/2003.08536) motivates behavioral novelty and selective transfer on top of [POET's paired environment/agent approach](https://arxiv.org/abs/1901.01753); our representation, learners, budgets and admission rules are much smaller.

The image-steering/accreditation connection remains measurement before retention: retain original input, declare a target and preservation region, record the predicted effect before acting, measure the actual result, and preserve unsuccessful attempts. v0.9 adds comparative compute costs and explicitly separate perceptual judgments. Exports can be inspected alongside the existing research work without treating numerical fit as accreditation of general ability.
