Why this journey?
The planner chooses a short journey by imagining what comes afterward. Here we test its alternatives from the same moment, and watch what changes when real care resumes.
Recorded observer comparisons on six previously used seeds. These alternatives reveal what happened in the simulator; they do not give the figure knowledge of its future.
Watch the continuations ↓Which choice earned its place?
Loading the fixed journey audit…
Each score chooses one journey. We then measure its 48-tick care cost after ordinary planning resumes. “Lowest cost” includes ties. The best alternative is found afterward, within the tested menu and window.
Forecast scores include an estimated reserve penalty for later care. Measured cost covers the 48 observed ticks only. A ranking difference can also reflect that terminal estimate. Lower care cost and more time with both beds viable are separate outcomes.
Watch the assumption separate from care.
The full initial journey is committed in both branches. After it ends, one follows the forecast’s recorded script; the other uses local readings and returns to ordinary planning.
The first seed’s first steady and compound checkpoints are shown regardless of outcome. Playback adds no simulated actions. Use the table to inspect every alternative, including options the planner rejected.
Does one model make a better choice?
The candidate keeps learning and switching among all its models, but scores each journey with the currently selected model alone. These are complete seasons on twelve fresh seeds, after freezing the change.
Loading the full-care comparison…
The first fresh seed’s steady and compound seasons are retained regardless of outcome. More economical computation is reported separately; the frozen promotion rule requires a care benefit.
Change the next question with evidence.
A retrospective better choice is a diagnostic lead. A controller must select it using only available observations and succeed in fresh complete seasons before it earns promotion.