# Two places, one memory — Conservatory 0.4

The home page is a new two-bed experiment. The single-bed nursery remains at `/nursery.html`; its original studies and physics remain unchanged. The Maze stays at `/ecology.html`. Nursery and Maze link to one another.

## What movement changes

Every cardinal grid move, inspection, actuator operation, wait, or blocked attempt consumes one physics tick. Both beds continue to gain or lose moisture during travel. A mechanism can be operated only from its station. Breadth-first routing around the two beds is authored, and the controller knows station coordinates. This is consequential travel, not learned navigation or physical manipulation.

The fern bed's viable interval is 35–65%; the reed bed's is 45–75%. Their evaporation rates, water gains and delivery lags differ. Valves share a reservoir with capacity 2.4, initial supply 1.9, and limited replenishment. Limited water is split proportionally to demand. Delivery takes 2–4 ticks, depending on the bed and configuration; water already in a pipe can arrive after a valve closes. Shade halves evaporation. A drain removes 0.06 moisture before that tick's weather and deliveries.

Controls sit directly below the environment. An intervention queues travel, then performs the operation, leaving automatic care paused. Valve buttons change opening by 50 percentage points. Visit A/B travels to the reading station. Pause movement cancels the remaining route; completed travel ticks still count. New season keeps learned response parameters and clears spatial observations. Reload replays and verifies the saved actions, then leaves the journey paused.

## What Lumen receives

`observeCare` supplies tick, position, current weather, moisture within one Manhattan step of a bed's sensor, equipment states within one step of each station, and reservoir level only near the tank. It excludes distant moisture, hidden response coefficients, the configuration seed, and delivery queues. Observation and evaluator traces are separate fields.

For each bed the controller remembers its last reading, reading time, known equipment state and an estimated present moisture. Estimates integrate observed weather and remembered valve/shade actions. Three simple response models consider delays of 0, 2 or 4 ticks; local returns update coefficients and model error. The smallest observed model error selects the current estimate. The displayed age is exact; the internal uncertainty increment is a heuristic, not calibrated confidence. Hidden delivery lag need not equal a model's candidate lag.

Action priorities, target commitment, station locations, thresholds, and pathfinding are authored. The controller learns limited response parameters. It does not learn its body or its motivation. The reservoir gauge, optional actual-moisture display and historical plots are observer instruments and never change the policy input. Scene colors normally use estimates; enabling actual moisture changes the renderer only.

## Frozen care comparison

Protocol `two-bed-comparison-1`: four training seasons and eight evaluation seasons per control, 240 actions each. Each control receives 2,880 actions including training; total 14,400. Evaluation resets spatial observations and starts from that control's training snapshot. The learned arm can adapt during an evaluation, but no evaluation updates the next evaluation's starting model.

| Control | Both beds viable | Mean water | Mean travel ticks |
| --- | ---: | ---: | ---: |
| Remember + learn | 52.60% | 3.468 | 175.9 |
| Remember + fixed model | 31.41% | 3.456 | 181.8 |
| Last reading only | 52.50% | 3.468 | 172.6 |
| Forget remote readings | 17.08% | 3.468 | 130.6 |
| Fixed 24-tick patrol | 24.95% | 3.468 | 72.5 |

“Both viable” is the fraction of ticks when both beds fall within their own intervals. This stricter metric differs from mean individual-bed viability. These eight evaluation seasons are descriptive evidence, not a statistical superiority claim. Remembering the last reading performs almost identically to fitting a model. The learned arm's 0.10 percentage-point edge does not establish a useful learning advantage. Most controls consume essentially all available supply.

The forgetting control clears observations, history and visits each tick but retains a committed route until arrival; otherwise it could not finish a journey. Both forgetting and patrol controls constrain which bed is considered using the same fixed alternating schedule. Consequently that contrast changes scheduling as well as memory and cannot isolate a pure causal memory effect. Fixed-model and last-reading controls are stronger comparisons for model usefulness.

Report: `experiments/results/care-study-cb6d3a66-c612-4838-bd27-5601a34a2be7.json`. Frozen sources and protocol: `experiments/frozen/care-v4/`. The former single-bed scores are not directly comparable because the physics, goals and observation boundary differ.

## Maze transfer comparison

The finite search borrows environment–learner pairing, mutation of eligible environments and transfer of solutions from [POET (Wang et al., 2019)](https://arxiv.org/abs/1901.01753). This is a small authored maze grammar with a tabular learner, not a reproduction of POET's bipedal system or a demonstration of unbounded novelty.

Protocol `maze-transfer-comparison-2` fixes three training seeds and eight fresh test genomes. For each training seed, evolution with transfer, evolution without cross-pair transfer, and independently generated random challenges each receive exactly 12,000 training/search actions. Selection, adaptation, transfer validation and training all count. Any remaining allocation trains existing pairs, including a final partial episode. The random control trains four independent memories round-robin on randomly drawn full-mechanism layouts. It therefore tests this complete curriculum alternative, not an isolated one-variable ablation of generation.

Every final population has four learners. All four face all eight test layouts with learned parameters frozen. Spatial memory is rebuilt from local observations within each test. No best learner is selected using test performance. A geometry-key exclusion check covers the full training archive and proposed candidates, ignoring seed-only differences. It found zero overlaps. There are 96 evaluation episodes per arm, clustered within three training seeds, not 96 independent training replicates.

| Curriculum | Test layouts solved | Mean attempts |
| --- | ---: | ---: |
| Evolution + transfer | 96/96 | 67.906 |
| Evolution without transfer | 96/96 | 66.396 |
| Random challenges | 96/96 | 67.063 |

Transfer is slightly slower here. All controls saturate success, so this study supplies no evidence of a transfer advantage. Internal accepted transfers do not override this end-to-end result. Total measured comparison cost is 127,331 actions: 108,000 training/search and 19,331 evaluation.

Report: `experiments/results/maze-study-a138d028-8b94-40bb-8c62-0d1ab01ddcc8.json`. Protocol and source bytes were frozen before evaluation in `experiments/frozen/maze-comparison-v2/`. Their manifest contains test genomes and SHA-256 hashes.

The first comparison stopped after a population-size assertion: reaching the budget between child admissions skipped oldest-pair retirement, leaving five learners. Its report and source snapshot remain available. The fix applies the same oldest-retirement rule even on partial generations. No test-performance selection was introduced. The corrected comparison uses fresh test layouts. The old attempt cost 14,140 completed actions plus a conservative 19,680-action reservation for the interrupted chunk; that reservation is an upper bound, not a measured action count. Both sets of test layouts are now exposed and must not be described as new validation on rerun.

## Runs on this Mac

The Maze's Evolving challenges panel now uses the local Node service. Switching pages or closing the browser does not stop it. A normal development run stops at 12 generations, 120,000 actions, 120 seconds of active worker time, or three generations with no admissions or accepted improvements. The plateau rule is heuristic, not proof that nothing more could be learned. The new seed-2718 run and its exact costs are retained in `experiments/results/v4-compute-audit.json`.

One lab worker runs at a time; the shared service dispatcher allows at most two concurrent compute tasks across the existing and new workers. Care and maze comparison jobs have 600-second active-time limits and action ceilings of 14,400 and 177,120 respectively. The latter reserves the maximum evaluation cost and can complete below that ceiling.

Private `data/lab.json` records completed chunks atomically. Pause prevents new dispatch; a current chunk may finish. Stop terminates current work and charges its full outstanding reservation. On service restart, unfinished runs pause and interrupted reservations are conservatively charged; Resume continues from the last completed chunk. Changed engine hashes block resumption. Mac sleep or stopping the service stops computation. There is no boot-time daemon, remote machine or automatic recurring run.

API: `GET /api/lab`, `POST /api/lab/jobs` with kind `maze`, `care-study` or `maze-study`, `GET /api/lab/:id`, and `POST /api/lab/:id/pause|resume|stop`. These use the existing same-origin and loopback protections. Full exports contain protocol, costs, rows and learned state. New runs never overwrite old jobs.

```sh
npm start
npm test
npm run test:service
npm run test:lab
```

Use the page's comparison buttons or Evolving challenges panel to queue a bounded run. Archived source modules expose deterministic creation/chunk functions for reproduction. Such reproduction uses exposed seeds. Service reports exclude development tests, manual browser journeys and training-only debug replays; the compute audit explicitly scopes its totals. Worker elapsed time is not a whole-machine energy measurement.

The connection to image steering remains an observation–intervention–prediction–evaluation discipline. These experiments do not establish better image editing, general intelligence or sentience.

### Trace timing

The frozen study records `observation` before an action, `next` and `actual` after the physics tick, and `beliefs` after remembering the action but before assimilating the next observation. Use those observation ticks when inspecting estimates; they are not simultaneous truth/estimate samples. The live UI assimilates the next observation immediately for its displayed chart. This does not change the next action within a season, but its end-of-season memory includes the final reading; the frozen study's training checkpoint precedes that final assimilation. The protocols and source snapshots preserve this distinction.
