# The Conservatory 0.3 — a bounded regulation experiment

The nursery asks an ongoing question: can a controller keep moisture between 35% and 70% as rain, evaporation, and valve response change? Each season contains 192 actions, with a hidden parameter change at step 96. It is a deliberately small synthetic environment, not a biological model or a test of consciousness.

The live page uses an online least-squares dynamics model and a six-step beam-search planner. The model learns from observed transitions. Planning, the target range, action vocabulary, model form, and initial coefficient estimates are authored. The initial model is already a useful prior; a fresh model is not an agent without knowledge.

## Observations and actions

The structured learner receives exactly five local readings: moisture, valve position, shade state, current rain, and heat. It does not receive seed, future weather, shift time, true dynamics coefficients, or the enclosing simulator. Actions are hold, open valve one increment, close valve one increment, toggle shade, and drain. Actions take one simulator step; they are actuator commands, not navigation movements.

The pixel arm receives only a 64×64 RGB software-camera view of the local instrument panel, plus ordinary reinforcement-learning reward and episode flags. It receives no tile labels or structured sensor vector. This is a raster observation with stable geometry, not a natural-image or first-person 3D benchmark. The large Three.js glasshouse is an observer visualization and is never passed to the learner. Palette, illumination, and small camera offsets can vary without changing physics.

`NurseryEnv` provides reset/step semantics in JavaScript. The Gymnasium adapter keeps the same Node simulator as the authoritative physics through a private subprocess pipe. Both pixel and vector adapters pass Gymnasium's environment checker. Evaluator reports are separate from policy observations.

## Frozen comparison

`experiments/frozen/regulation-v1/manifest.json` was written before the first regulation study. It fixes eight training seeds, eight evaluation seeds, four perturbations on two evaluation seeds, ten controls, 192 actions per episode, and prediction horizons 1, 4, and 12. These seeds are now exposed. Reruns are replications, not new held-out evidence.

Each control consumes 4,608 actual environment actions: 1,536 training, 1,536 evaluation, and 1,536 perturbation actions. Fresh memory receives the same training allocation and discards its learned state before each evaluation episode. Fixed and memory-ablated controls also consume the training allocation. Evaluation episodes start independently from the trained snapshot, rather than accumulating evaluation experience across seeds. Structured learners adapt within each episode.

Controls: retained memory; fresh memory; one-step planning; memory erased after each action; reduced sensors; valve actions removed; an authored threshold rule; all unique archived sequences; no sequence archive; and retained memory with an explicitly privileged guard. Their observation/action/assistance differences are intentional ablations, not interchangeable performance scores.

The report includes total actions, all-phase viability, water expenditure, model-evaluation count, worker elapsed time, and paired differences against fresh memory. Training and perturbation costs are included. Equal interaction budgets do not mean equal CPU budgets. The approximate paired interval uses the eight evaluation seeds and is descriptive; it does not establish generalization across other environment families.

Common probes use one prescribed 192-step reference trajectory, regenerated for each of ten models: 1,920 additional simulator transitions and 2,720 predicted rollout steps per complete study. These costs are separate from the 46,080 learner actions, giving 48,000 actual simulator transitions overall. Every frozen model predicts the same action sequences and moisture outcomes, without learning from them. One, four, and twelve-step RMSE is consequently comparable across models; on-policy errors alone are not. Reduced sensing also changes the information available at probe starts. There are 16 starts per horizon, not independent replications of the training process. Worker episode elapsed times exclude probe execution and service overhead.

Perturbations: initial moisture ±0.001; two-step observation delay; and two-step actuator lag. Recovery is the number of post-shift steps until eight consecutive viable readings; eight can mean the controller never left the viable range. A missing value means the criterion was not reached. This definition is not time from a detected failure.

## Reusable behavior and ancestry

The controller archives distinct six-action sequences from training, with start/end moisture, water cost, source seed, and model parent. A sequence is marked demonstrated when it reduces distance to the moisture target by more than 0.025 without guard assistance. The useful-sequence arm offers only these sequences to the planner; the all-sequence arm offers every distinct stored sequence. Options compete by predicted rollout cost and are reconsidered every action. A demonstrated training effect does not certify that a sequence transfers.

The archive has a 32-sequence cap. This is a finite behavior library, not an evolutionary population. Every training checkpoint has an ID and parent in the study's lineage list. Downloads retain the final model, sequence evidence, and intermediate ancestry metadata.

The separate challenge generator makes four parameter mutations of a parent environment in each of four generations. Selection admits intermediate difficulty (20–95% viability), then prefers the candidate nearest 65%. The continuing model inherits only selected learning. Version 1 lost all four previously demonstrated capabilities. Version 2 adds a separate five-case regression gate; all four proposed promotions were rejected. Four of five reporting capabilities remained demonstrated and none were acquired. The gate and reporting cases are distinct; the reporting suite is already exposed development evidence. This is a measured plateau, not open-ended evolution or a reproduction of POET.

## The neural baseline

The local experiment uses the author-maintained [DreamerV3 implementation](https://github.com/danijar/dreamerv3), pinned to commit `e3f02248693a79dc8b0ebd62c93683888ddaccfe`. CPU-only dependencies live in `.venv-dreamer`, with exact installed versions recorded in `experiments/dreamer-requirements.lock.txt`. No system Python packages were changed.

The run uses `size1m` (686,327 optimizer parameters in this observation setup), batch size 4, sequence length 16, 15 imagined steps, and one update per 16 interactions after replay warm-up. It completed 1,536 training actions and 93 gradient updates in 45.4 seconds including compilation and evaluation. A 900-second wall limit bounded the process. The debug configuration was not used.

Eight frozen evaluation seeds yielded 31.7% viable time for trained weights, 39.2% for untrained weights, and 32.2% for trained weights under changed appearance. Trained weights and policy RNG counters are restored before each paired appearance condition. No learning benefit is established by this small, single-training-seed run. A small appearance difference does not establish visual invariance.

Neural evaluation freezes weights, while the structured controls adapt within episodes. Observation representation, model size, optimization, and compute differ. These results are a bounded feasibility baseline, not an algorithm ranking. The neural model runs through the CLI and its evidence is visible in the page; it does not drive the live glasshouse controls. Local checkpoints and complete evaluation traces are in `experiments/dreamer-runs/pixels-frozen-v1/`.

## Constraints and claims

The optional guard inspects privileged simulator state and may replace an action whose one-step predicted moisture falls outside [0.18, 0.90]. Proposed action, executed action, and assistance are logged. The guard is not credited as learner competence. It is a discrete intervention and does not guarantee viable moisture, especially under lag or uncontrollable disturbances. It is not an implementation of a continuous control-barrier-function theorem.

Verified claims are finite: tested observation boundaries, replayed physics, bounded accounting, interface contracts, and restart recovery. Observed claims concern the recorded runs. AGI, sentience, unbounded novelty, and universal safety remain unestablished. More machines or runtime alone does not establish any of those properties.

The first regulation report recorded full sensor readings in the reduced-sensor trace even though its controller correctly received masked readings. The final boundary rerun corrects this evidence field; its physical outcomes and trained models are compared in `regulation-boundary-audit.json`. Both versions are retained. Dreamer's baseline used full-sensor rendering and is unaffected by that trace-only correction.

## What the sources contribute

| Source | Concrete use and limit |
| --- | --- |
| [Turing (1936)](https://www.cs.virginia.edu/~robins/Turing_Paper_1936.pdf), [Rice (1953)](https://doi.org/10.1090/S0002-9947-1953-0053041-6), [Hamkins & Nenu](https://arxiv.org/abs/2407.00680v3) | Scope verification to finite checks. Undecidability of general semantic questions does not make these specific bounded tests impossible; distinguish Turing's original printing/circle-free arguments from the modern halting formulation. |
| [Wolfram (1985)](https://content.wolfram.com/sw-publications/2020/07/undecidability-intractability-theoretical-physics.pdf) | Report computation and prediction horizons. No assertion that all finite prediction requires full simulation. |
| [Lorenz (1963)](https://people.ucsc.edu/~rmont/classes/chao/2013/Orig_Papers/Lorentz_EN.pdf) | Paired initial-condition and disturbance experiments. They do not prove that this nursery is chaotic. |
| [Ashby (1956), chapter 11](https://ashby.info/Ashby-Introduction-to-Cybernetics.pdf) | Ongoing regulation, with sensor and action ablations. Do not turn requisite variety into an unqualified equation between sensor count and intelligence. |
| [Ames et al. (2019)](https://arxiv.org/abs/1903.11199) | Explicit state constraints and auditable intervention. Formal forward-invariance results require model and feasibility assumptions this toy guard does not supply. |
| [Lenski et al. (2003)](https://www.nature.com/articles/nature01568) | Track useful intermediate behaviors, their ancestry, and transfer cost. The original digital-organism findings are not a sentience recipe. |
| [Taylor et al. (2016)](https://doi.org/10.1162/ARTL_A_00210) | Distinguish adaptive novelty, capability retention, combinations, and plateaus from more random room seeds. |
| [ECMA-262](https://tc39.es/ecma262/multipage/overview.html), [WHATWG workers](https://html.spec.whatwg.org/multipage/workers.html) | Separate deterministic simulator logic from host execution and lifecycle. The Node service owns durable studies; browser rendering may suspend. |
| [WebRTC recommendation (2025)](https://www.w3.org/TR/2025/REC-webrtc-20250313/) | Reserved for a future explicitly requested multi-machine experience-sharing experiment. The present deployment is loopback-only on this Mac; WebRTC would not supply persistence or knowledge merging by itself. |
| [POET](https://arxiv.org/abs/1901.01753) | Motivates parameter variation, a minimal selection criterion, and transfer checks. Our finite curriculum omits its larger coevolving population and transfer machinery. |

The [image-steering connection](research-link.md) remains observation → intervention → measured consequence → qualified claim, with negative results retained as evidence.
