## Direct raster trial added in v0.7

The link now includes actual image-domain evidence: `/image.html` applies measured local/global raster controls to procedural RGB images, scores the selected region and surroundings separately, and records predictions, measured errors, acceptance and undo. [Method](image-steering-v1.md) and [completed findings](v7-findings.md) distinguish numerical preservation from perceptual quality or accreditation. The research repository was read without modification.

# From image steering to an observable learning system

The Conservatory is a separate executable research instrument. The source library remains at `[local research source]`; this implementation does not alter its findings or claim an improvement to its editing results. The first table describes the v0.2 object-room experiment. The v0.3 nursery connection follows below.

| Library input | Concrete implementation | Boundary |
| --- | --- | --- |
| `references/30_control_observation.md` | `sense()` supplies at most 13 nearby tile observations. The controller receives a remembered map; rendering has a different, full-world view. | Sensors provide semantic classes and absolute position. There is no learned visual perception. |
| `references/35_identifiability_experimental_design.md` | Predictions are captured before updating on each movement or interaction. Sample counts and observed local tile transitions accompany predictions. | A 0.5 prior is not evidence. Empirical frequencies are not calibrated uncertainty intervals. The model does not identify every hidden mechanism. |
| `references/65_spatial_fields_dependent_prediction.md` | An action changes a spatial scene. Crates move, shutters control growth, valves clear a crossing. A local sensor records visible consequences and spatial memory. | These dynamics are authored. The learner predicts one-step action effects, not general spatial semantics or delayed causal structure. |
| `references/70_decision_policy_comparison.md` | Retain every seed, failure, training attempt, evaluation attempt and model lineage. Compare a model-backed learner with sample-only values, fresh memory and an exploration control. | Lower prediction error does not entail better goal-directed behavior. Mandatory evidence logging is not itself a superiority result. |

The 16-room v1 comparison records this distinction directly: retained memory improves its on-policy object-change Brier score (0.142 versus 0.260), but takes more attempts than fresh memory (80.25 versus 72.19). Both succeed in 16/16 generated transfer rooms. The scored actions differ between policies, so even the prediction-error comparison is not a common held-out prediction set. Sample-only learning scores 0.140 and averages 79.31 attempts. The model-backed method does not beat that baseline here. The exploration control reaches 12/16 and averages 341.63 attempts, counting all failures at the 420-attempt limit.

Training costs are 973 attempts for the retained model and 984 for sample-only values. All 96 episodes together use 11,131 observed attempts. Saved evidence is `experiments/results/transfer-v1.json`. These seeds become exposed development data after this run. No parameter adjustment followed by another run on the same seeds should be presented as fresh validation.

## Recommendations recorded after v0.2

1. **Credit assignment:** distinguish useful consequences from merely repeatable object motion. Repeated crate pushing can earn local novelty while delaying the resource. The current short-horizon reward and state abstraction are candidate causes, not established causal findings.
2. **Common prediction probes:** freeze actions and observation histories before comparing model error. Report both probe accuracy and end-to-end task outcomes.
3. **Pixel observation adapter:** render a first-person or local orthographic view, remove tile labels and coordinates, and preserve the same simulator/action/reward contracts. Evaluate representation learning separately from control.
4. **Established world-model baseline:** integrate the author-maintained [DreamerV3 implementation](https://github.com/danijar/dreamerv3) through its environment interface; pin a commit and dependencies, demonstrate tiny-run compatibility, and report full training compute. It is a distinct recurrent model and actor-critic baseline. It had not been installed or evaluated in v0.2; v0.3 completes a bounded local run.
5. **Open-ended environment search:** the [POET paper](https://arxiv.org/abs/1901.01753) supplies a relevant environment-agent coevolution precedent. This release uses a fixed generator and fixed training families. It does not implement POET's novelty/minimal-criterion loop or population transfer selection.

The artistic inputs remain [Computational Beings](https://un1crom.github.io/worksonbecoming/computational-beings/learn/) and [On Noah's Ark There Is No Theater](https://un1crom.github.io/worksonbecoming/on-noahs-ark-there-is-no-theater/). They inform the invitation, glasshouse atmosphere and participatory observation; they do not validate the learning claims.

## Version 0.3: decisions have to survive measurement

The nursery extends the same loop to ongoing regulation. Five local sensors feed a learned dynamics model; a six-step planner chooses valve, shade, or drain actions; the record pairs the prediction with its observed consequence and cost. A separate RGB-only camera adapter feeds a recurrent DreamerV3 baseline. Its visual representation is a local raster instrument panel, not a general scene understanding result.

| Library question | Nursery evidence |
| --- | --- |
| What could the controller observe? (30) | Explicit sensor boundary, delayed/reduced sensing controls, raw camera view, and observation records distinct from privileged evaluator state. |
| Can the model identify a useful consequence? (35) | Prescribed common action probes at horizons 1/4/12, without fitting; separate on-policy outcomes and prediction errors. |
| Does an intervention remain useful after conditions change? (65) | Weather/actuator changes, initial-state perturbations, sensor/actuator lag, synchronized histories, and retained action sequences. |
| Does the decision policy improve outcomes for its cost? (70) | Matched interaction allocations including training; separate compute and water costs; guard assistance exposed; negative baseline and curriculum results retained. |

The conclusions remain narrow. Retained and fresh models both achieve 100% viable time in the frozen nursery, so the task does not distinguish a memory advantage. The short pixel Dreamer run performs worse than its untrained control. Harder selected environments initially erase old capabilities; a regression gate preserves the earlier model and reports a plateau. These are useful failed hypotheses rather than claims of intelligence growth.

In the earlier rooms, adding observed push clearance and direction reduces mean attempts from 80.25 to 72.125 on the already exposed seeds. The reward-only change yields 78.9375. This supports testing the context hypothesis on a new, frozen navigation study; it does not retroactively turn the old comparison into a positive transfer result.

The practical connection to image steering is an auditable evaluation pattern: specify available observations, define an intervention, predict before observing, compare outcomes under a fixed allocation, and keep failures. A claim about better image editing still requires a new image-editing capability and direct evaluation in that domain. See [the nursery protocol and sources](regulation.md) for full scope and reproducible commands.

## Version 0.4: observations expire while the body travels

The two-bed experiment turns spatial absence into a measurable cost: every movement advances the weather, remote readings age, and the controller must estimate delayed deliveries from its own observation/action history. Shared water couples otherwise distinct care tasks. The observer can reveal hidden moisture without changing the learner's observations. This sharpens the library's control/observation distinction (30) and its identifiability question (35).

The memory/model comparison (70) gives every control the same total action allocation and retains both success and resource use. Learned response fitting reaches 52.60% both-bed viable time, essentially tied with last-reading memory at 52.50%. Lower results for the forgetting control also reflect its scheduling constraint; they do not isolate a pure causal memory benefit. The result supports inspecting what the model contributes before claiming that added model complexity helps.

The spatial-generation comparison (65/70) evaluates all four members of each final maze population on eight layouts excluded from training and selection, across three training seeds. Transfer, no transfer and random curricula all solve 96/96 episodes. Transfer averages 67.91 attempts versus 66.40 without transfer and 67.06 with random challenges. Internal transfer improvements have not produced a held-out task advantage.

An interrupted first comparison remains in the compute record. Its population-cap bug was fixed under a new frozen source/protocol version and evaluated on fresh layouts. This is the same evidence discipline that image-editing accreditation needs: preserve failed runs, state what changed, separate observation from privileged evaluation, and measure the proposed capability in its actual domain. There is still no new image-editing experiment here. Full sources, control definitions, accounting and limitations are in [the v0.4 protocol](care-v4.md).


## Version 0.5: prediction, adaptation, and the cost of a wrong explanation

The new common-action histories address the confound in earlier model-error comparisons: all predictors receive the same observations and actions. At twelve ticks, adaptive moisture prediction has 0.80 percentage points mean absolute error versus 2.64 for frozen models and 7.62 for holding the last reading. End-to-end viability is measured separately; its adaptive-minus-frozen paired interval still includes zero.

The supplied RFS material adds the question of continuity under shared consequences. A pipe fault and a reservoir shortage can look similar locally but call for different responses. Physical resource exhaustion can persist after a fault ends, and model updates made during scarcity can also damage later performance. These are observation/identification issues (30/35), followed by a policy comparison (70), rather than evidence that a longer-running process has become sentient.

The maze adds a finite PAIRED-inspired performance-gap search and a retention gate. One selected room improved on two further exploration seeds, with all four retention rooms still solved. Two individual retention rooms became slower, despite improved aggregate cost. Those details remain in the record.

For image steering, the transferable deliverable is a testable sequence: preserve the earlier model, detect a prediction failure, compare plausible causes using controlled interventions, and require improvement on the target plus retention elsewhere before adopting a revision. Applying that sequence to image edits still requires direct image-domain implementation and measurement. [RFS reading](rfs-reading.md); [v0.5 protocol and evidence](adaptation-v5.md).


## Version 0.6: protected memory does not guarantee a correct explanation

A crossed-component follow-up isolates the old fresh-state regression in retained bed estimates. Restoring those estimates after supply shortages recovers 87.50% viability, while restoring the supply estimate alone leaves 53.28%. v0.6 therefore protects the normal checkpoint and tests context selection, fitting, investigations and remaining-season resource costs separately.

The new twelve-habitat comparison preserves normal-checkpoint performance but does not establish an overall care advantage. Fixed candidates average more viable time than fitted candidates, and the rejection rule flags none of 24 unfamiliar compound-fault windows. This gives the accreditation work a concrete distinction: preserving a known capability (65/70) is different from correctly identifying why a new intervention failed (30/35). A plausible label must not be mistaken for a verified cause.

For a later image-domain test, preserve the baseline edit response, make a localized requested edit, record candidate predictions before the action, measure the target and untouched regions separately, and permit rejection of every available explanation. The novelty-rejection failure here argues for testing that last condition directly. There is still no new image-editing capability or image-domain result in this release. See [the frozen protocol and all outcomes](context-v6.md).
