# Conservatory research roadmap

Updated 20 September 2026, after the completed v0.19 full-season executed-work comparison.

This is a proposed sequence with evidence gates, not a promise that each stage will succeed or a schedule for AGI. The roadmap began after v0.9; reading-value, timing, model/planning, prediction and follow-up diagnostics are complete in v0.10–v0.19. Earlier evidence remains preserved.

## Direction

Build a visible, persistent learner whose accumulated experience helps it adapt to unfamiliar conditions, choose useful investigations, preserve earlier abilities and act within a finite budget. Make its behavior and the evidence behind our claims inspectable.

The next milestone is **reliable improvement from experience**. Longer runs and additional machines become useful when we can demonstrate what additional experience buys. Work stays on this Mac for now.

## Starting evidence

The [v0.9 findings](docs/v9-findings.md) establish the starting point:

- Nursery: learned investigation selection did not improve care over ensemble control. The difference was −0.65 percentage points, with a paired 95% interval of −1.64 to +0.35. Training labels combine the physical benefit of an intervention with the benefit of its readings.
- Maze: fresh learning succeeded in 48/48 rooms; repaired transferred priors succeeded in 45/48. All three failures combined changed appearance with distant wiring and exhausted search budgets. The curriculum recorded no winning transfers.
- Images: coarse alignment preserved 60/60 successes with 95.18% fewer predicted pixel-channel values. This is a specific computational saving, not a measurement of total runtime or energy.
- Photographic evidence remains narrow: three correlated crops from one source, with no human ratings collected.

Existing replay, lineage, retention checks, bounded jobs, exports and earlier pixel-learning baselines are foundations to extend. Their existence is not a new research result.

## Stage 1 — Establish when investigation pays

**Status: Stage 1 remains active. Reading-value, timing, model/planning, weather-care and journey-scoring comparisons are complete; reliable learned improvement remains unproved.**

The [completed diagnostic](docs/v10-findings.md) found useful reading-dependent tests at 28/180 fresh checkpoints (15.56%; 95% interval 9.48–21.64%). Mean retrospectively attainable gain was 0.85 percentage points (0.16–1.54 pp). Both intervals cross the declared targets, so the frozen decision is inconclusive. No revised selector was trained or promoted. The live nursery now exposes synchronized with/without-reading comparisons and the first action difference.

The [v0.11 timed-care comparison](docs/v11-findings.md) completed 300 full seasons. Timely care averaged 84.92% both-bed viability, versus 83.91% for ensemble care. The paired gain was +1.01 pp (95% seed-cluster interval -0.14 pp to +2.17 pp). The timely policy did not meet the frozen promotion rule. Ensemble care remains the default. Timely care recorded 76 early stops and 211 saved scheduled ticks. Those counts alone do not prove better care. The optional policy and synchronized recorded journeys are inspectable in the nursery. All earlier evidence is preserved.

The [v0.12 diagnostic](docs/v12-findings.md) compared five branches at 120 exposed checkpoints. Supplying known current coefficients improved care by +1.46 pp at the existing 48-tick horizon (95% descriptive interval −0.72 to +3.63 pp). Extending fitted-model planning to 96 ticks changed care by −0.43 pp (−1.41 to +0.55 pp) while nearly doubling branch model transitions. Adding full current state beyond known coefficients changed care by +0.07 pp (−0.36 to +0.50 pp). All intervals include zero, and benefits vary across seeds. This makes prediction accuracy a useful next investigation; it does not prove that inaccurate parameters are the main bottleneck. No controller was promoted. The new visual replay makes these observer comparisons inspectable.

The [focused code review](docs/v12-review.md) also fixed saved-journey cancellation history, export pausing, save-queue recovery and late completion beyond the time budget. Those fixes leave earlier scientific source and evidence unchanged. 102 tests, both HTTP suites and independent replay checks passed.

The [v0.13 prediction audit and care comparison](docs/v13-findings.md) ruled out the suspected calibration problem on the exposed panel. Recorded future weather removed much of the long-range forecast error. A forecaster fitted only to observed weather history lowered development prediction error, but its fresh full-care result was +0.30 pp versus ensemble care (95% interval −1.57 to +2.17 pp). It did not meet the promotion rule. The new prediction page makes the conditional forecasts and both complete care journeys inspectable. All 47 frozen scientific sources and prior evidence are preserved.

The [v0.14 journey audit and fresh comparison](docs/v14-findings.md) tested all twelve ordinary journeys at 60 exposed checkpoints. The existing score selected a lowest-cost continuation at 25/60. Scoring with the active model alone reduced severe short-window errors, mainly in one seed, but failed the fresh complete-season comparison: **−2.25 pp versus ensemble care (95% interval −4.19 to −0.30)**. Ensemble care remains the default. All assumed follow-up scripts differed from subsequent real planning, but their effects interact with prediction errors and terminal valuation. The new journey page exposes these comparisons and the full-season evidence. 115 tests, both HTTP suites and exact retained replays passed.

The [v0.15 planning trace and pilot](docs/v15-findings.md) found a repeated promise to close bed B four ticks later while actual B moisture crossed its viable limit. All 74 sampled plans used differing follow-up actions across models. One common-script candidate repaired the traced continuation, but failed complete development seasons: **66.71% versus 82.86%, −16.15 pp** across three exposed seeds and five conditions. It was retired without consuming fresh cases. The trace page exposes the mismatch, local repair and broader failure together. 119 tests, both HTTP suites and retained replay checks passed.

The [v0.16 two-decision pilot](docs/v16-findings.md) grouped predicted local reading histories and required each group to share forecast actions. It passed the information-boundary tests but failed the fixed development gate: **71.60% versus 82.86%, −11.26 pp**, with ten of fifteen candidate seasons reaching the original computation reserve. The retained compound example was already worse before the limit. The candidate is retired; no fresh cases were consumed. The observation-branch page exposes forecast groups, actual next decisions, complete seasons and budget stops. 124 tests, both HTTP suites and exact retained replays passed.

The [v0.17 terminal-value diagnostic](docs/v17-findings.md) found that removing the reserve penalty changed **0/60 choices**. An explicit 96-tick forecast changed **35/60**, but slightly reduced mean 96-tick viability (**−0.07 pp**) and did not earn a full-season pilot. All 720 original 48-tick comparisons reproduced exactly. A separate two-case profile found **67.63% and 78.31% repeated model calls**, mostly in lookahead projections. This is an exact reuse opportunity, not an achieved speedup. No fresh cases were consumed; ensemble remains the live default. All 58 frozen scientific sources and prior evidence are preserved.

The [v0.18 exact-reuse comparison](docs/v18-findings.md) shares identical waiting projections within each tail evaluation. Search time fell **17.18% and 21.27%** in two exposed benchmarks, with **25.87% and 27.85% fewer executed model transitions**. Scores, four short trajectories, two complete seasons and budget stops remained exact; logical budget charges were preserved. At four fixed replanning forks, withholding later moisture/tank readings left every proposed journey unchanged. Committing the original forecast's next operation repaired the exposed failure but harmed two successful controls. All 61 scientific sources and earlier evidence are preserved; the live ensemble remains unchanged.

The [v0.19 executed-work comparison](docs/v19-findings.md) ran 45 complete seasons under the original ceilings. Charging only executed model transitions improved the shared two-decision planner from **71.60% to 76.25%** (+4.65 pp), reducing budget-limited seasons from ten to one. Ensemble still achieved **82.86%**, so the candidate failed the development gate by **6.61 pp** and received no fresh evaluation or promotion. Every pre-limit trajectory matched exactly; all old control rows reproduced. In the retained compound case, avoiding the tick-437 stop changed actions but did not improve care. All 63 scientific sources and prior evidence are preserved.

**Next bounded work:** Freeze an audit of the exposed pre-budget replanning forks. Compare the promised next operation with ordinary alternatives from the same current fitted beliefs, using a common continuation rule and matched physical execution. Retain actual local readings and the v0.18 controls where commitment harmed care. Determine whether the changed estimate or the way future actions are scored makes the old promise lose. Only then propose a bounded rule for retaining or replacing a follow-up, and require the complete-season gate before fresh evaluation. Keep exact reuse and actual-work accounting available for offline experiments; more computation alone did not repair this policy.

Primary question: When does acquiring a reading change a later decision enough to improve care after paying for travel, time and the intervention itself?

1. Use development checkpoints to compare an identical physical intervention with and without its readings, plus ordinary care. Separate the reading contribution from the intervention contribution. Keep privileged simulator state confined to offline diagnostics.
2. Estimate whether there is enough actionable information benefit to learn. The existing retrospective checkpoints suggest limited opportunities; do not assume a larger training set will solve that.
3. If headroom exists, train a cost-aware choice among tests and ordinary care. Record prediction uncertainty and allow abstention. Use the separated effects to explain a choice while evaluating its complete downstream outcome.
4. Freeze the method, training data and decision criteria before drawing a fresh evaluation set. Compare against ensemble care and existing investigation baselines under matched action and computation ceilings; report training cost separately as well as total cost.
5. In replay, show which reading changed the plan, what it cost and what happened. Label retrospective comparisons as diagnostics; do not present a hidden-state comparison as the learner's foresight.

**Advance if:** fresh paired evaluation supports a practically meaningful care benefit under the declared budget, or a separately preregistered noninferiority comparison supports equivalent care at meaningfully lower cost. Choose the primary claim, practical margin and sample plan before evaluation; do not switch claims after seeing results.

**Change direction if:** readings rarely improve attainable outcomes, or the selector still fails on fresh evaluation. Retain ensemble care and investigate the sensing, model family or planning horizon that limits improvement. More test starts, lower prediction error and more training examples do not by themselves satisfy this gate.

Follow-on work after this primary question:

- Diagnose the three exposed maze failures as development cases. Test confidence-based reuse, physical revalidation and a reserved recovery budget on new rooms. Fresh learning remains the comparison; these three known failures cannot serve as new validation.
- Expand photographic sources and collect independent blinded judgments of intended change and outside preservation. Numerical success and perceptual quality remain separate outcomes.

## Stage 2 — Learn beyond the authored menu

**Status: conditional on Stage 1 diagnostics.**

The present learner selects within supplied model families and investigation actions. Expand what it can represent: relationships between actions and effects, delays, dependencies and combinations of mechanisms. Change appearance and physical rules independently so that superficial recognition cannot masquerade as causal transfer.

The experimental separation is informed by [CausalWorld](https://arxiv.org/abs/2010.04296), which uses interventions to define controlled training and evaluation distributions. Applying that principle to these habitats is our proposed adaptation, not a reproduction of its robotics benchmark.

**Advance if:** the learner adapts to held-out mechanism combinations without adding a special rule for each evaluation case, within a fixed budget and without privileged observations. Include failure cases and compare with the authored baseline.

**Change direction if:** an observation cannot distinguish candidate mechanisms, or the representation cannot express the needed relationship. Diagnose that limitation before increasing model size. Revisit the earlier pixel baseline only with a specific perception bottleneck and a suitable budget.

## Stage 3 — Demonstrate cumulative learning

**Status: later.**

Run sequences of changing habitats with retained memory, tested against reset learners. Measure which knowledge transfers, when it should be discarded and whether old abilities survive. Extend the existing curriculum and retention machinery rather than treating an expanding archive as evidence of learning.

[Enhanced POET](https://arxiv.org/abs/2003.08536) motivates behavioral novelty and selective transfer. Our bounded curriculum is an adaptation; it does not establish open-ended evolution.

**Advance if:** multiple independent learner histories show increasing held-out task coverage or decreasing adaptation cost while retaining prior abilities. Count unsuccessful transfers, training and challenge-generation costs. Reserve evaluation families outside curriculum selection.

**Change direction if:** challenge diversity is mostly cosmetic, reuse fails to outperform relearning, or new skills erase earlier ones. Inspect representation and memory policy before extending run duration.

## Stage 4 — Make the research useful outside its habitat

**Status: later, with photographic review preparation available now.**

Strengthen the connection to `russell/image-steering-accreditation` through a common evidence contract: observations, predictions recorded before action, interventions, costs, protected regions, outcomes and reproducible checkpoints. Share evidence structure before assuming that learned parameters should transfer across nursery, maze and images.

Use a varied photographic benchmark, blinded human judgments and an independently replayable evidence package. Keep simulator diagnostics distinct from live image observations. See the existing [research connection](docs/research-link.md).

**Advance if:** the practical image task improves on a declared comparator, preservation remains acceptable and someone can reproduce the numerical findings from the exported package. Human review must support any perceptual claim.

## The visual experience develops throughout

The figure should have a visible history that matters: where it has been, what it observed, which expectation failed and why its next action changed. Put that history beside the world and its intervention controls. Use actual recorded states and decisions; distinguish measured facts, model estimates and observer explanations.

Each research stage should make at least one consequential behavior easier to see. The artistic direction remains a participatory learning habitat, with its uncertainty and dependence on care visible. A more expressive presentation alone is not stronger evidence of intelligence.

## Scale and the long-term ambition

Broader generality remains an ambition, not a scheduled release. Consider additional machines only after a bounded comparison shows that additional experience or computation improves held-out capability. Keep checkpointing, lineage and complete cost accounting at the larger scale.

Learning, general intelligence and consciousness require different evidence. We have no result showing that running this project longer would produce sentience. [Butlin and colleagues](https://arxiv.org/abs/2308.08708) propose assessing AI consciousness through indicators derived from scientific theories; that approach does not supply a runtime threshold or turn task success into proof of consciousness.

## How turns hand off

This file is the project map; versioned method documents and frozen reports hold the scientific detail. A turn need not finish a stage or create a new version.

For each substantive work cycle, record:

1. The primary question and why the existing evidence makes it worth asking.
2. The bounded change or experiment, its budget and success criteria.
3. What changed in the usable application and what verification completed.
4. The measured result, uncertainty, cost and limitations.
5. A decision: retain, revise or retire the approach; then the next unresolved question.

Preserve negative results and exposed test cases. Update this map when evidence changes the route. Leave interrupted work with a concrete checkpoint and remaining steps. Completion of a turn is a handoff in the project, not a reset of its direction.
