# Adaptation, with a record of failure

19 September 2026. Version `adaptation-0.5.0`; nursery comparison `adaptation-comparison-1`; maze search `maze-weakness-search-1`.

The nursery now compares predicted journeys, fits delayed response models from local observations, and interrupts a journey after persistent prediction errors. It retains earlier models and tests unannounced local and shared disturbances. This is a small model-based control experiment with authored equations, semantic sensors, and a finite action grammar. It does not establish AGI, sentience, or open-ended learning.

## What changed

The initial development planner scored a journey followed by waiting. This often undervalued closing valves and drawing shades, wasted water, and failed despite a much stronger full-information reference controller. The final planner evaluates each possible first journey over 64 simulated ticks, using a feedback controller in the forecast state for its continuation. It commits the first journey and replans at its completion or when a mismatch interrupts it. The displayed alternatives are the actual computed forecasts, including water, travel, end moisture, and cost. Continuation is authored; its forecasts use the learner's estimated state and parameters, not hidden simulator values.

The learner fits evaporation and water delivery for delay hypotheses of one through six ticks using up to 32 observed intervals per bed. Intervals involving predicted supply limitation or clipped moisture are excluded from fitting. Earlier models can be reconsidered after a mismatch. The source rate is estimated when the body visits the tank. Local observations, known actions, and current weather are its inputs. Unknown moisture, physical coefficients, the fault label and its schedule are unavailable to the ordinary policy. Future weather is held constant in forecasts.

Two consecutive errors above an age-adjusted threshold can trigger a mismatch event, with a 24-tick cooldown. This is an uncalibrated heuristic: an alarm is not a diagnosis, and silence does not establish that nothing changed. The inspection bonus favors stale readings and permits an eight-tick local investigation after an alarm. Manual interventions clear the pending automatic plan. Browser checkpoints replay both actions and post-observation model updates.

## Frozen comparison

Source and protocol were frozen before evaluating the specified configurations in `experiments/frozen/adaptation-v5/manifest.json`. Nine source files have individual SHA-256 hashes. Development used habitat seeds 7201, 7213 and 7237. The formal comparison uses five independent base seeds: 31011, 31033, 31057, 31079 and 31103.

Each base habitat gets two adaptive training seasons. Every evaluation starts from a clone of the same trained model for that habitat, empty local readings, and a fresh physical reset. Training is computed once and shared across controls; it is not counted six times. The two evaluation seeds vary initial moisture, disturbance location and onset. Physical base parameters and weather phase remain familiar. This tests adaptation within trained habitats, not transfer to unfamiliar physical laws.

Each 384-tick evaluation has one condition:

- Steady: no additional disturbance, but the ordinary weather continues to vary.
- Pipe: one bed's delivery gain becomes 30% of its original value.
- Drying: one bed's evaporation coefficient becomes 180% of its original value.
- Supply: shared tank refill becomes 20% of its original value.

Disturbance begins at a seeded tick between 112 and 135 and ends at tick 288. Returning the coefficient to normal does not refill the tank, reset moisture, or erase water in transit.

| Control | Definition | Both beds viable |
| --- | --- | ---: |
| Adaptive planner | Online fitting, mismatch interruption, inspection bonus | 78.87% |
| Frozen model | Trained parameters held fixed; readings still assimilated; no mismatch interruption | 67.23% |
| No inspection bonus | Online fitting and mismatch interruption; ordinary readings still occur; extended investigation/tank-inspection options omitted | 76.48% |
| Last reading | Holds remote moisture estimates between visits; no fitting or mismatch interruption; the planner still uses the trained dynamics for hypothetical futures | 56.87% |
| Full-state planner | Same journey planner with true current moisture, pipes, tank and effective physical coefficients; no future weather or fault schedule | 88.27% |
| Full-state rule | Authored valve/shade rule with full current state and coefficients, under the same travel and operation costs | 82.27% |

There are 40 evaluation episodes per control. The full-state references have privileged information; neither is an optimal bound. Adaptive versus frozen changes both fitting and mismatch-triggered replanning. The comparison does not isolate either component alone. The last-reading control is a temporal-estimation ablation, not a policy without any model.

Adaptive minus frozen averages **11.64 percentage points**, with five paired habitat differences of 25.13, 5.14, 3.35, 1.60 and 22.98 points. A paired t interval with four degrees of freedom is **−2.55 to +25.83 points**. All five observed differences favor adaptation, but the interval includes zero. Five heterogeneous habitats do not establish a reliable general advantage. No tuning followed inspection of these evaluation results.

| Condition | Adaptive during change | Frozen during change | Adaptive after return |
| --- | ---: | ---: | ---: |
| Steady, equivalent windows | 91.75% | 77.00% | 74.38% |
| Pipe fault | 90.92% | 52.97% | 67.60% |
| Increased drying | 80.08% | 62.86% | 20.00% |
| Shared supply shortage | 84.40% | 71.69% | 33.96% |

Adaptive alarms occur during 5/10 pipe, 3/10 drying, and 6/10 supply disturbances. Mean detection delays among detected episodes are 21, 101.7, and 108.7 ticks, respectively. These conditional means exclude non-detections and must not be presented as universal response times. The export retains null delays. Recovery means twelve consecutive viable ticks after the first post-change loss; zero means viability was maintained during the entire disturbance, and null means no qualifying recovery was observed. Later relapse remains possible. Alarms after the return are counted separately from pre-change/steady false alarms.

## Prediction on common histories

Twenty prescribed trajectories (five habitats × four conditions) cycle physical valve, shade, and inspection actions. Adaptive, frozen, and last-reading predictors receive identical prefix histories. At each forecast origin they see the same current weather and prescribed future actions, never future observations. All future trajectories are evaluated against the same recorded physical outcomes.

| Horizon | Adaptive MAE | Frozen MAE | Held-last-reading MAE |
| --- | ---: | ---: | ---: |
| 1 tick | 0.320 points | 1.206 points | 3.872 points |
| 6 ticks | 0.507 points | 1.942 points | 5.885 points |
| 12 ticks | 0.803 points | 2.638 points | 7.618 points |

These are moisture percentage points, averaged over both beds. Unseen-bed scores and counts are also exported. Predictions within a trajectory are correlated, so the many scored ticks are not independent replicates. The common-history result supports better prediction in these prescribed contexts; it does not by itself prove improved control or an accurate diagnosis of a fault.

## Why returning to normal did not restore performance

A separate **post-hoc diagnostic on exposed seeds** reproduces the first evaluation seed in each of the twenty habitat/condition combinations. It then compares pre-evaluation and post-evaluation model snapshots in identical fresh, steady, 128-tick habitats, with learning disabled. The reset supplies the same full tank, initial moisture and empty local readings to both models. It changes physical state and horizon, so its results cannot be directly compared as an isolated tank intervention against the original 384-tick run.

| Previous condition | Fresh habitat, before model | Fresh habitat, after model |
| --- | ---: | ---: |
| Steady | 87.34% | 81.72% |
| Pipe | 87.34% | 82.66% |
| Drying | 87.34% | 76.88% |
| Shared supply | 87.34% | 53.28% |

Both physical resource exhaustion and degraded retained models matter. In all five sampled drying and supply episodes, the original run ends with an empty tank. Even after a fresh reset, the model retained after a supply shortage performs substantially worse than the pre-evaluation model. We have not isolated whether that loss comes from stale supply estimates, misidentified bed coefficients, or their interaction.

An optimistic resource bound assumes both shades drawn from tick zero, no tank overflow, no waste, and no undelivered water at the end. Three of these twenty cases still require more water than could be supplied: two drying cases and one supply case. Full-duration viability is impossible in those cases even under these generous assumptions. A zero deficit is inconclusive; it does not prove that a feasible route or controller exists.

This narrows the next research question: distinguish shared supply restriction from local delivery failure, protect a useful pre-change model, and test restoration in matched physical states. Longer uninterrupted running alone would not resolve that identification problem.

## Maze weakness search

The finite search proposes twelve rooms per round for three rounds. A witness gets two training episodes on each candidate; both witness and learner are then evaluated with learning frozen. A solved witness and a loss gap equivalent to at least 9.6 attempts make a candidate eligible. Two further exploration seeds must preserve its advantage. Learner adaptation is tested on two additional seeds. Four regression rooms must retain success and total attempts must stay within 5% of the prior total. Individual rooms can become slower under that aggregate gate.

The completed search used **11,665 actions**, proposed 36 rooms, found four eligible gaps, and retained one improvement. On its two adaptation checks, attempts fell **70→60** and **82→72**. All four retention rooms remained solved, with total attempts **352→337**. Two individual retention rooms worsened (62→72 and 113→122), which is retained in the report. Rounds two and three found no eligible new gap. These selection and retention rooms participate in development; this is not held-out maze generalization or evidence of unbounded novelty.

The maze's Evolving challenges panel includes this search, pause/resume/stop controls, export, candidate records, and a replay of the successful witness. Nursery/Maze navigation remains reciprocal.

## RFS and primary research

The supplied RFS position paper contributes a useful distinction between unpredictable behavior and demonstrated adaptation. Its emphasis on continuity, shared consequences, and measured recovery motivated the separate local/supply disturbances and the physical-state-versus-model diagnostic. The kit's body-dependent sensing and action costs are promising future variables, but they were not silently imported as learning capabilities. See [the RFS reading](rfs-reading.md).

- [PETS — Chua et al., 2018](https://arxiv.org/abs/1805.12114): learned dynamics used in predictive control. This implementation has small deterministic fitted models and bounded journey forecasts; it does not implement PETS's probabilistic ensembles or trajectory sampling.
- [Bayesian Online Changepoint Detection — Adams and MacKay, 2007](https://arxiv.org/abs/0710.3742): separates inference about change from ordinary parameter estimation. Our threshold-and-cooldown detector is not Bayesian change-point inference and produces no calibrated posterior.
- [PAIRED — Dennis et al., 2020](https://arxiv.org/abs/2012.02096): environment generation using a performance gap with a successful counterpart motivates the maze search. Our finite candidate grammar, affordance learner, and validation gates differ from adversarial reinforcement learning.
- [POET — Wang et al., 2019](https://arxiv.org/abs/1901.01753): the preserved population-based maze search remains separately available.

## Reproduction and accounting

Start the loopback service with `npm start`; open `/` for v0.5, `/care.html` for v0.4, `/nursery.html` for v0.3, and `/ecology.html` for the maze. New service runs can be started from the corresponding controls. Completed evaluations are now exposed; rerunning them is reproduction, not fresh validation.

- Main report: `experiments/results/adaptation-study-932e6796-f161-401e-8b4c-3ef613fb90a6.json` — 103,680 environment actions, 198,680,747 model transitions, 51.8 seconds of active worker time.
- Maze report: `experiments/results/maze-frontier-fbfbe253-eca2-4cf3-8956-a7ee972ff1e5.json` — 11,665 environment actions, 1.8 seconds of active worker time.
- Recovery diagnostic: `experiments/results/recovery-diagnosis-v1.json` and its pre-run manifest — 12,800 environment actions, 31,067,369 model transitions. Reproduce with `node experiments/recovery-diagnosis.mjs`; this script imports the frozen source.
- Audited evidence total: **128,145 environment actions**, **229,748,116 model transitions**, no uncertain service actions. These units are not interchangeable. Model transitions are deterministic forecast calculations, not learner interactions or FLOPs.
- `experiments/results/v5-compute-audit.json` also inventories retained development reports. Its evidence total excludes debugging, unit/integration tests, browser animation and manual journeys; not every early debugging run had a complete numerical log. Worker time excludes startup, rendering, machine energy and wall time outside a worker.

The service has no active research jobs after these bounded runs. No additional machine, external agent, physical hardware, or remote deployment was used. Earlier engines, frozen comparisons, and negative results remain intact.
