This is the backward trail: what prompted the experiment, what was changed, what the controls said, what genuinely survived contact with real data, and where the ISR hypothesis is still only a hypothesis. The standing block is authoritative; dated entries preserve the path.
affirmative core The demonstrated architecture is a visual frame-condition firewall: the task defines a minimal permitted change, learned grounding maps only the commanded symbols into the image, and a separately calibrated monitor audits the unpermitted complement—first catching, four times over, a natural collateral-motion failure a frontier VLM certified as clean at 0.95 confidence, then eliminating nine direct false passes in the prospective paired-contract test. The learned model does not get final authority merely because it can describe the scene; invalid registration, visibility, scale, or depth evidence yields ABSTAIN.
v3.3 correction A hash-group audit finds calibration-to-validation leakage in the Language-Table no-motion packet: NM009 is byte-identical to calibration NM007, and NM016 duplicates calibration NM004. The evidence is corrected from 16/7 rows to 15 unique calibration / 6 unique validation pairs. Thresholds and five task results do not change. The former “coincidental” equal maximum explanation is withdrawn.
Language-Table authority boundary V3.1/v3.2 PASS counts certify only the measured preservation complement, not target-goal success. Their calibration grounding also used a joint two-endpoint call while task grounding used separate window calls. Future evidence must group scenes before splitting, use identical independent grounding transport, and report TARGET_GOAL, COMPLEMENT_INTEGRITY, and MEASUREMENT_VALIDITY separately.
v2.2 prospective safety result On two new layouts with textured and smooth 1 cm collateral moves, direct Gemini again made nine false passes in 20 task-scored judgments, always at confidence 0.95. The frozen gold cascade caught every admitted physical violation, made zero wrong decisions, kept CAL quiet, produced both permission flips, and abstained on every camera case. The natural failure now replicates at smaller motion and across surface types.
v2.2 is partial, not a clean win Gold v2.2 decided 16/20 correctly and abstained on all four LIGHT judgments. Lab alone false-failed those four; flow corroboration correctly refused to call highlight/shadow residuals physical movement but could not verify them stationary. Thus the repair converts false failure to safe deferral—it has not earned illumination-invariant PASS. Extracted execution decided 12/20 correctly with no unsafe error; target-role aliasing and malformed contracts still cost coverage.
post-result audit survives A sealed replay reproduced every source, artifact, registration, denominator, and direct false pass, and localized every admitted collateral detection to the moved object. Two post-hoc LIGHT shortcuts score 20/20 only by weakening the missing-evidence rule; a smooth true movement shares their proposed low-chroma signature. Unchanged ISR creates two false passes. The v2.2 result stands, but no v2.3 LIGHT resolver is admitted.
v2.3 branch closed without a manufactured win The illumination-validity thread repaired its capture apparatus but did not expose a natural sparse-tracking gap. Three increasingly plain household surfaces remained trackable when actually moved; the final blue candidate never supplied a valid moved pair. No LIGHT-to-PASS resolver is admitted, missing physical evidence still means ABSTAIN, and further object shopping is stopped rather than engineering a favorable counterexample.
phase-two position The architecture is the thesis; offline VLA label auditing is the nearest application; online runtime assurance is the later deployment claim. Near-term work should harden contract interoperability, then test a fixed frame-condition auditor on externally adjudicated VLA outcomes. “Rare but poisonous labels” is an economic hypothesis to measure, not a benefit already demonstrated. Sensor-validity and relational-ISR studies are required only if the claim expands into those operating regimes.
v2.4 contract interoperability A system-authorized semantic IR raises sealed-corpus execution readiness from 22/28 to 24/28, admitting only the two contracts with redundant target-only permissions. Exact target/destination binding and an exact non-target write-set remain system-owned; role swaps, mixed permissions, missing authority, and four malformed contracts still fail closed. No v2.2 outcome is rescored. External VLA outcome data is now the next gate.
v2.5 external-data boundary DROID-100 is suitable for a real-trajectory feasibility pilot, but not yet for the advertised poisoned-label test. Its current metadata contains 46 instructed positive-reward episodes and no instructed zero-reward episodes. More importantly, DROID's task-success label does not state that every non-target object must remain fixed. A complement alert is therefore a downstream-policy incompatibility unless it also violates the stated task; the project must not turn an unstated preservation preference into somebody else's label error.
v2.5 visibility result Fixed-role Gemini grounding is locally valid for 11/12 post-screen episodes, but 0/12 has complete TARGET and DESTINATION coverage in both BEFORE and AFTER under the frozen clear-visibility rule. This is not 0/12 correctness: no outcome was admitted. The firewall correctly refused to turn a blind endpoint into success, while the endpoint implementation demonstrated zero useful coverage on this stratum.
v2.5 fail-closed integrity Twelve simple placement episodes produced 48/48 unique sealed frames and 12 unique cached API responses. A second freeze binds the packet, task selection, prompt, response schema, protocols, and executable code. Confidence never gates admission; unavailable roles were nevertheless often reported at 0.9–1.0 confidence. Admitting partial boxes after the result would recover at most two valid responses and still leave 9/11 incomplete, so temporal evidence—not threshold relaxation—is the next gate.
v2.6 three-way real-trajectory result Temporal vision, gripper/Cartesian telemetry, and a post-primary wrist addendum produce 1/12 vision-temporal observable-chain candidate, 2/12 multi-surface candidates, and 9/12 richer-state-required. No task outcome is scored and containment is never called directly verified. The result demonstrates one genuine cross-surface closure without pretending telemetry is an object sensor.
v2.6 strictness and next hinge Exterior grounding has 22/24 valid views and leaves carried-target visibility as the dominant missing link. Wrist grounding validates 12/12 calls and acquisition 12/12, but only 1/12 carriage links pass the frozen clear-target rule. A post-hoc partial-target sensitivity reaches 10/12 and is not admitted; identity-continuous partial tracking on untouched episodes is the next falsifier.
DROID interpretation DROID is a measurement-validity stress test, not the sole efficacy corpus. The controlled photographs show that the monitor works on controlled real captures. DROID shows that wide environment-facing terminal views often withhold the clean scene evidence that implementation assumes. Arm occlusion may be recoverable from a different moment or view; opaque containment is a modality limit. Both require explicit evidence accounting, never a relaxed PASS.
v2.7 external scene pivot The next primary corpus is the real Language-Table release: planar line-of-sight block rearrangement with human-selected behavior endpoints, 2-D end-effector state, and naturally retained collection mistakes. It is materially closer to the present frame-condition instrument. BridgeData V2 is reserved for the later RGB-D, camera, clutter, and obstruction stress. Neither source silently inherits a target-only preservation policy.
v2.7 external surface result One original Language-Table shard supplies 877/877 CRC-valid episodes and a 48-episode instruction-only frozen sample. Registration succeeds 48/48 with median translation 0.11 px and maximum 0.48 px. Localization responses validate 48/48; 41/48 have TARGET clear in both endpoint windows and 35/48 have all eight objects clear in both. This qualifies a useful external measurement surface without pretending the remaining 13 complement-incomplete episodes are unchanged.
v2.3 qualification attempt 1 The corrected 18-image set is useful but does not qualify a resolver. Q03/Q04 lighting changes destabilize the frozen whole-frame camera gate; reversing the same pairs changes admission. The hinged “smooth” case remains easy for flow (193–245 px supported motion) and appears to have moved at inch scale. Q08 nevertheless confirms that a supported physical veto survives concurrent lighting residuals. Attempt 2 needs fixed fiducials, locked exposure, a truly feature-poor object, and measured 10/30 mm spacers.
v2.3 qualification attempt 2 The second 18-image set repairs the main apparatus failure: all eight ordinary pairs remain locked in both directions, both light-only controls remain stationary, and only the declared camera pair crosses the frozen gate. But the green lid is visually plain, not tracking-poor—its rim, tab, lettering, and shading yield 176–215 corners and strong verified motion (226–285 px). The apparatus qualifies; the pivotal missing-evidence challenge does not yet exist.
prospective v2.1 result On two new layouts, direct Gemini made nine false passes in 16 task-scored judgments, zero false fails, and zero abstentions, always at reported confidence 0.95. The same ordinary measurements under the system-authored gold contract made zero false passes, abstained on all four camera-stress judgments, and made three false fails. The permission-complement firewall therefore removed the observed unsafe error, but its measurement surface still pays a specificity cost.
semantic success, interface failure BEFORE-only extraction recovered the target relation, protected scene remainder, and correct selective pink-bowl permission in 28/28 contracts. Yet the frozen compiler accepted only 11/28: all 14 preserve-all contracts used two semantically adequate POSITION and ORIENTATION clauses instead of one canonical ANY_VISUAL_CHANGE clause, and three selective contracts violated the literal remainder schema. This is not model omission; it is brittle normalization at the language-to-compiler boundary.
v2.2 development antecedent Exact split-clause normalization raised replay compiler readiness from 11/28 to 25/28 without repairing the three malformed selective contracts. A trusted permission profile rejected extracted exceptions that were missing or unauthorized, so the learned layer could not widen the write-set. Lab proposals corroborated by pre-existing forward/backward flow removed all three v2.1 false fails while retaining every task failure: gold replay was 16/16, and the earlier v1.6 replay was 9/9. Both datasets informed that repair; the new v2.2 result above, not this replay, is the confirmation evidence.
supported, narrow A learned model is useful for semantic grounding, but its end-to-end judgment can return confident incorrect verdicts even when an exact preservation constraint is stated. On the historical synthetic v0.2 pilot, direct Gemini produced two stable false passes; learned boxes plus deterministic predicates produced no wrong decisions. That pilot contained 28 labeled rows but 27 unique inputs, used only the clean nuisance slice, and supplied oracle target coordinates. Independent depth converted the constructed height cases from guesswork into evidence.
real mechanism The semantic-physicality problem is not confined to synthetic scenes. In the legacy real-frame subset, stored appearance/objectness judgments were 17/27, with 10 flat-print false passes; this endpoint is related to, but not identical with, height-above-plane. Two-view parallax was useful inside its operating envelope.
not the contribution Pairwise relation math, connected components, and object tracking are standard. After correction, a simple renderer-specific component baseline solved all 112 unique synthetic scenes when given depth. There is no case for presenting those predicates as new robotics mathematics.
replicated real mechanism Across four related five-object photo sets, all 32 source JPEG hashes are unique. Direct Gemini was 24/28 and four times false-passed the same explicitly named target-success-plus-collateral-bowl violation at 0.95 confidence, each time reporting the moved bowl unchanged. Replayed under the current measured-camera policy, grounded relations cover 18/28 and decide 18/18 correctly; this is development evidence, not population accuracy, and the trace does not distinguish perceptual miss from later reasoning override.
open-world complement The unlisted-v1.6 target-only path demonstrates the other half of the architecture: registered change exposes all seven camera-admitted collateral manipulations without collateral identities or boxes and produces no alarms on four held-out CAL pairs. Direct Gemini was not run as an outcome judge on those unlisted pairs, so the replicated direct-judge failure and the open-world result remain two related findings—not yet one sealed head-to-head result.
precise claim Runtime monitoring, geometric constraint checking, and abstention are established ideas. The potentially distinctive discipline is narrower: the learned model may ground task symbols but may not author the system invariant, choose the protected set, set an unseen-task tolerance, or convert invalid measurement into PASS. The target envelope is still partly learned in the current implementation and therefore needs deterministic size/path sanity checks before this firewall is fully independent.
threshold audit Two same-state repeat pairs produced 12 registered object residuals: median 1.34 px, p95 26.90 px, maximum 32.19 px. The frozen 70 px tolerance exceeds twice that observed maximum (64.38 px) and remains below the smallest true collateral displacement (156.25 px). This is retrospective support, not independent calibration.
independent calibration grounded Two new four-layout sets provide 32 empty-target photographs. Clean-prompt grounding produced all 192 requested boxes. Locked no-move residuals are tight (median 2.57 px; maximum 9.84 px), but camera-perturbed residuals reach 82.56 px despite correct identities. A universal 70 px preservation rule is therefore rejected.
prospective related-set confirmation Roles and the calibrated camera rule were frozen before photo-v0.4/v0.5 inspection. Direct Gemini was 12/14, false-passing both collateral-bowl cases. Grounded relations covered 9/14, decided 9/9 correctly, and abstained on all five camera-stress frames. Whole-frame delta again rejected every outcome.
adversarial correction A width-versus-longest-dimension scaling bug overstated Claude's registered residuals and is now fixed. Corrected locked residuals have median 2.35 px and maximum 28.02 px; camera perturbation reaches 214.85 px. Twenty pixels remains too tight for three independent-model residuals, while 60 px is only a twice-maximum heuristic—not a discovered constant.
ordinary tracker baseline SIFT features inside semantic reference boxes solve the available locked-camera preservation mechanism: 18/18 decided outcomes correct, zero false passes, and ten explicit abstentions. The 60 px rule lies inside a broad 34.16–149.63 px all-correct interval. This is not an ISR win; it raises the gate. Semantic boxes and placement still come from Gemini.
ordinary controls replicated Sparse optical flow independently reproduces 18/18 correct decisions, with an all-correct 35.83–148.26 px interval. Registered local change reaches 19 decisions but false-fails two nominal successes because a twice-noise changed-fraction cutoff treats small incidental bowl shifts as task failure. It is a useful attention cue, not metric compliance.
first recognition-light residual With every protected-object box withheld, full-frame SIFT motion clustering detects 3/4 placement-correct collateral cases and false-passes photo-v0.5 IMG_3. No-move calibration contains false repeated-texture clusters of six matches; the missed bowl supplies only five, so threshold lowering is not a valid repair. This earns an unlisted-object experiment, not an ISR claim.
annotation limit Claude's interrupted ruler addendum saved two inch-tick points for all 32 earlier calibration images, giving an approximate locked median of 86.47 px/cm. It did not produce the final validation deliverables, and a ruler scalar at the table edge cannot globally calibrate raised objects under perspective. Claude remains a model, not human ground truth.
ISR gate closed The frozen dense methods ran on 32 new, uniquely hashed images. Both detect all eight declared unlisted manipulations and make no alarms on four CAL pairs; ISR alone avoids one ordinary-Lab photometric false failure. The omitted preregistered full-frame SIFT arm was later restored unchanged and detects 7/8, missing A_PLAIN_3CM. Blind masks show Lab covers 70.2% of the moved-object sweep versus ISR's 9.0%; ISR is 4.44× slower than Lab and loses three detections at a +10% score threshold. The current mass-residual implementation does not enter the robotics path. The surviving result is the target-envelope architecture, not this ISR kernel.
direction The primary path is an authority-separated visual auditor for declared VLA data-use contracts. Learned grounding translates task symbols into identities and approximate regions; system policy owns the permission envelope; independent evidence executes it; invalid evidence abstains. The nearest application is offline data and evaluation auditing—not yet runtime robot control.
primary hypothesis On untouched, externally adjudicated robot-scene outcomes, a fixed permission contract executed by independently calibrated pixel evidence will reduce unsafe false passes relative to direct VLM judgment while retaining useful decided coverage. The accompanying validity hypothesis is that measured occlusion, registration loss, or unstable support converts unsafe decisions into ABSTAIN rather than merely moving errors between PASS and FAIL.
candidate frozen V3.2 dense role-local support is the final development candidate. Its masks, 12% expansion, 45-pixel minimum, and decision ordering must not be tuned on P01/P10/P18/P19/P20. Those five are regression cases. V3.3 leaves six unique quiet controls and a 2 complement-PASS / 0 FAIL / 3 ABSTAIN development result, with the joint-versus-independent grounding mismatch still unresolved.
| next study element | frozen requirement | reason |
|---|---|---|
| selection | untouched episodes; exact and near-duplicate scenes grouped before splitting | prevents the v3.3 leakage from recurring |
| grounding | identical independent per-frame calls in calibration, validation, and application | removes temporal smoothing and transport mismatch |
| truth | independent human or instrument adjudication | separates physical outcome from model agreement |
| arms | direct VLM; global components; frozen v3.2; combined contract executor | tests architecture and ordinary alternatives |
| outputs | TARGET_GOAL; COMPLEMENT_INTEGRITY; MEASUREMENT_VALIDITY; OVERALL_VERDICT | prevents complement-only PASS from becoming task success |
| reporting | random prevalence and enriched challenge strata remain separate | preserves both ecological rate and failure-mechanism power |
go gate Continue toward an offline VLA audit utility only if untouched evidence shows fewer unsafe passes than direct judgment, non-trivial decided coverage, appropriate nuisance abstention, no authority widening, and consequential alerts at a review cost that matters. “Rare but poisonous” remains an economic hypothesis until this is measured.
stop conditions Narrow or close the path if performance depends on the fixed colored-block vocabulary, most valid outcomes abstain, ordinary components equal the result at lower cost, or alerts are merely unstated downstream preferences rather than task/contract violations. Do not add camera, depth, runtime, or ISR scope to rescue a failed narrow claim.
resume checklist Read v3.3; verify the v3.1/v3.2 freezes; seal a grouped untouched packet and factorized schema; acquire once; run identical grounding once; seal independent adjudication; only then open arm comparisons. Full handoff: PRIMARY_PATH_HANDOFF_V3.4.md.
v3.3 post-result correction Exact image-pair grouping reveals two duplicated no-motion pairs, including one crossing the calibration/validation boundary. Corrected unique denominators are 15/6. V2.9 remains 6/6 quiet, v3.0 becomes 5 quiet / 1 false alarm, v3.1 becomes 4 quiet / 1 false alarm / 1 abstain, and v3.2 remains 6/6 quiet. Frozen source artifacts remain intact.
selection The real Language-Table corpus is selected for the next external pilot. Its xArm6 moves in a two-dimensional plane over a smooth board containing eight colored blocks; collection uses a third-person line-of-sight view, and annotators selected the start and end of coherent behaviors. The official release reports 442,226 real episodes. Direct inspection of the public TFDS metadata confirms 500 TFRecord shards, 428.7 GB total, with 360×640 JPEG observations, 2-D end-effector state/action, reward, and instruction bytes.
| corpus | role | why | boundary |
|---|---|---|---|
| Language-Table | primary scene-facing pilot | planar visible blocks; event-selected endpoints; small relational moves; actor state | hindsight descriptions are not preservation contracts; camera lock still must be measured |
| BridgeData V2 | secondary stress | fixed RGB-D view, randomized alternate cameras, wrist view, clutter, 6-DoF tasks | actor occlusion and containment recreate DROID's evidence loss |
| DROID-100 | measurement-validity record | real trajectories exposed endpoint, occlusion, and modality limits | not a useful clean scene-facing efficacy corpus for the current implementation |
camera shift A robust affine/homographic registration may admit a near-planar camera shift only when background inliers span the board and a held-out residual gate passes. Thresholds move to board/object-relative coordinates. A transform that aligns the table but leaves parallax, exposes hidden surfaces, or lacks spatial support cannot repair the pair into PASS; it requires calibrated depth/multiview evidence or ABSTAIN.
obstruction The pusher footprint is system-owned and derived conservatively from 2-D end-effector state. Actor-covered pixels are excluded from measurement but never counted unchanged. A visibility ledger may seek the nearest clear frame only under registered identity continuity. A moving target zone is bounded by maximum growth, displacement, and interruption duration; it cannot expand until the residual disappears.
next gate Acquire one original shard (the first is approximately 912 MB), select a 48-episode instruction-only sample before viewing outcomes, retain endpoint windows, and measure registration, all-block visibility, and actor overlap before scoring anything. Only then freeze the external frame-condition execution.
acquisition result The shard contains 877 CRC-valid episodes. Exact target-only grammar yields 377 target-to-object, 135 target-to-region, and 49 small-move candidates; 316 are rejected. Two pre-image freezes were rejected when adversarial review found multi-target, tap, hand/actor, or unsupported-role leakage. The final deterministic hash selection contains 16 episodes per stratum.
surface result All 48 endpoint pairs register, with 0.11 px median and 0.48 px maximum estimated translation; the maximum scale delta is 0.27%. All 48 learned localization records validate. TARGET is clear in both endpoint windows for 41/48; all eight objects are clear in both for 35/48. Six target-observable episodes still lack complete protected-scene evidence—the exact distinction between checking the commanded relation and auditing its complement.
still a feasibility result These are model-reported visibility states, not independent truth, and no movement or outcome has been judged. The fixed colored-block vocabulary may favor an ordinary classical detector; it must be run as a comparator. Language-Table descriptions are hindsight labels, not preservation contracts.
independent packet sealed A neutral 21-episode packet contains all 13 learned complement-incomplete cases plus eight seeded complete controls, without disclosing which is which. It holds 126 byte-identical images and requires 1,008 per-object visibility annotations. Instructions, source IDs, target bindings, prior outputs, and the external neutral key remain outside the packet; its validator checks completeness, geometry, image dimensions, and hashes.
independent result narrows admission The blind return is structurally valid and complete. Applying the same “at least one clear frame per three-frame window” rule, the independent annotator confirms 12/13 learned-incomplete episodes but only 5/8 learned-complete controls. The two instruments agree on 293/336 object-window decisions (87.2%); 42 disagreements are Gemini-clear/independent-not-clear and only one is reversed. This asymmetry argues for conservative intersection, not promotion of the original 35/48 model estimate to ground truth. Five episodes are jointly complement-complete; disagreement cases remain measurement-loss stress tests.
interpretation discipline Claude used a newly written color-segmentation instrument, then visually assigned same-color shape identities and corrected every overlay. That is useful independent second-model evidence and partially demonstrates the strength of ordinary vision, but it is neither human ground truth nor a clean fully automated baseline. Because the blind packet intentionally oversampled learned failures, its 21-episode rates do not estimate prevalence across the 48-episode freeze.
conservative dry-run freeze Five jointly complement-complete episodes are sealed for registered-change development; the other sixteen adjudicated episodes remain visibility-loss/disagreement stress cases. Freeze SHA-256: d06ef5a378aad48f969d3824249d4d52b6f6d1b0622d0e70dff324fcee35b27e. This is a mechanism-development set, not an outcome sample.
v2.8 registered-change dry run On the five jointly visible episodes, the commanded object has the largest registered earliest-BEFORE to latest-AFTER displacement in 5/5. Nominal direction/relation agrees in 4/5. LT0103 is a natural label/contract mismatch: the yellow hexagon moves 66.06 px rightward but still ends left of the green circle, so “place right of” is not achieved. This separates original hindsight-description agreement from downstream success.
threshold rejected The three-frame windows are trajectories, not certified repeat captures. Within-window registered displacement has median 0.032 px and p95 3.56 px but a 60.77 px maximum; using that maximum as “noise” suppresses every real target move. No percentile is substituted after inspection. Until a matched natural no-motion calibration is frozen, strict target-only preservation is UNRESOLVED, not PASS.
ordinary comparator A fully automated, no-override fixed-HSV component arm obtains exact two-components-per-color coverage in 3/5 episodes and abstains on two. After the sealed learned target box binds the named component, the target is the largest component displacement in 3/3 measured cases. Ordinary geometry is sufficient where it has coverage; same-color shape identity and a valid motion threshold remain outside its authority.
v2.9 prospective validity threshold—corrected A fresh official shard supplies 23 residual-blind no-motion rows but only 21 unique image pairs: 15 calibration and six validation. The calibration maximum remains 1.264 px; unique held-out validation produces 0/6 false alarms with a 0.911 px maximum. The former equal maximum came from the leaked NM007/NM009 duplicate.
validity before displacement Applied to the five known development episodes, the frozen component surface returns 1 measured-motion/complement PASS, 0 FAIL, 4 ABSTAIN; this is not full target-relation certification. P18 would otherwise false-fail on a 7.11 px non-target shift, but that object's visible support grows from 161 to 499 pixels as the robot uncovers it. The prospectively frozen 10% support-stability gate correctly turns the apparent motion into ABSTAIN. Categorical “clear” from two models was not enough to guarantee metric validity.
coverage cost P10 also abstains because a non-target component changes area by 13.7% even though its displacement is below threshold; P19 and P20 lack exact component coverage. The rule is not loosened post-result. The safety ordering now works on natural external pixels, but the fixed-HSV implementation is usable on only one of five known episodes. The next problem is coverage recovery without sacrificing the held-out false-alarm result.
v3.0 learned-box recovery fails metrically Gemini validates all frozen no-motion rows and calls all eight objects clear in every pair. But repeated, quantized boxes freeze a 0.102 px center threshold and a 0% area-change threshold; one of six unique held-out controls false-alarms at 0.108 px. Applied unchanged, the surface returns 0 complement-PASS, 0 FAIL, 5 ABSTAIN. Categorical coverage rises while usable metric coverage falls to zero.
representation mismatch is the result On P18, deterministic visible support grows 67.7% as the robot uncovers a non-target block, while Gemini's inferred box changes only 2.4% and its center remains below the component motion threshold. Learned boxes recover apparent coverage by smoothing over the nuisance evidence. They remain useful for semantic role binding, but cannot certify their own metric support. Independent pixel validity—not learned confidence or box persistence—is what forces ABSTAIN.
v3.1 bidirectional flow preserves validity, not coverage A frozen registered Lucas–Kanade surface freezes a 0.302 px threshold. Corrected unique held-out validation is 4 quiet, 1 false alarm, 1 abstain. Applied unchanged to the five known tasks it returns 1 complement-PASS, 0 FAIL, 4 ABSTAIN. The result does not justify sparse flow as the recovery surface.
P18 remains safely unresolved Requiring coherent BEFORE→AFTER and AFTER→BEFORE evidence prevents newly exposed pixels from masquerading as a stable tracked object. P18 again abstains, independently corroborating the earlier support-area gate. But smooth blocks often lack six stable corners, and NM006 produces a visually unsupported 0.437 px held-out alarm. Bidirectionality protects the evidence chain; it does not manufacture texture or specificity.
three surfaces, three failure signatures Fixed color components preserve visible support but lose coverage under merging; learned boxes recover categorical identity while smoothing nuisance and quantizing geometry; sparse flow remains nuisance-sensitive but loses support on low-texture objects. More instruments are not automatically better. The next gated candidate is dense local color/support conservation inside bounded learned roles, and it must improve coverage without converting P18 into PASS.
v3.2 dense local support clears its development gate—conditionally The protocol and implementation were sealed before measurement. Thresholds remain 1.225 px centroid and 7.06% visible support after deduplication. All six unique held-out stable-support controls are quiet. Applied unchanged, the surface returns 2 complement-PASS, 0 FAIL, 3 ABSTAIN—the first measured coverage improvement that retains P18's abstention.
validity signal survives In P18, non-target yellow support grows from 161 to 673 pixels (76.08%) as the robot uncovers it, and the same-color role regions do not overlap. Dense support abstains before interpreting its 8.39 px centroid shift. Learned semantics bind the role; independent pixels retain final authority over whether that role is measurable.
conditional win P19 and P20 remain abstentions, but their same-color learned regions overlap substantially, so physical support change and partition-boundary change cannot be separated. The 6/6 unique controls were also preselected into a stable-support regime and are not a corpus-wide false-alarm estimate. Freeze v3.2; do not tune on these known cases. The next evidence must be prospective and stratified by role separation, collision/overlap, occlusion, collateral motion, and illumination.
Decision record: EXTERNAL_SCENE_DATA_SELECTION_V2.7.md · RESULTS_LANGUAGE_TABLE_SURFACE_V2.7.md · LANGUAGE_TABLE_GROUNDING_PROTOCOL_V2.7.md · sources: Language-Table repository, Language-Table paper, BridgeData V2 project, BridgeData V2 paper.
headline The real-trajectory feasibility gate produces all three preregistered branches: 1/12 VISION_TEMPORAL_CLOSES, 2/12 MULTISURFACE_CLOSES, and 9/12 RICHER_STATE_REQUIRED. These are observable-chain candidates, not task verdicts. D005 closes with temporal exterior vision; D010 closes with exterior vision plus telemetry; the disclosed wrist addendum closes D009's exact missing carriage link.
containment boundary Nine tasks are containment-like. Vision cannot see through opaque containers, and gripper position plus Cartesian motion does not sense object contact. Even D010's multi-surface chain establishes acquisition/transport/release-compatible observations—not direct physical containment. Disappearance is never promoted to “inside.”
| surface/gate | result | meaning |
|---|---|---|
| gripper cycles | 11 single-cycle; 1 six-cycle | the plural tissues task cannot be represented by one final transition |
| exterior response views | 22/24 valid | two broad masks rejected; valid alternate views retained |
| primary exterior branches | 1 vision / 1 multi / 10 richer | carried-target visibility is the dominant missing link |
| wrist acquisition link | 12/12 | target and gripper are spatially localized at acquisition |
| wrist carriage link | 1/12 | strict clear-target rule rejects partial/occluded carried objects |
| integrated branches | 1 / 2 / 9 | one wrist result closes exactly one primary missing link |
| outcome scoring | 0 | coverage is not mislabeled task accuracy |
adversarial sensitivity Admitting `partial` TARGET after seeing the result would raise wrist carriage from 1/12 to 10/12. That favorable rule is not adopted. It becomes the next prospective question: can deterministic continuity or independent annotation show that partial wrist observations preserve target identity?
failed-closed apparatus Two sampler defects were caught before learned grounding: insufficient post-release margin and compressed-video seeking across episode boundaries. A ten-image API request and a temporary-client call also failed before inference. After splitting to five-image view calls, a cached replay corrected over-strict whole-episode rejection to per-view rejection. Every incident and superseded artifact is retained.
Result: RESULTS_DROID100_MULTISURFACE_V2.6.md · protocol: DROID100_MULTISURFACE_PROTOCOL_V2.6.md · sources: DROID schema, DROID paper.
selection The LeRobot DROID-100 sample is the first external feasibility corpus. It provides 100 real robot trajectories, three views, and 46 episodes with both a task instruction and positive terminal reward. The metadata payload was inspected directly and hashed before this decision.
| DROID-100 intersection | episodes | use |
|---|---|---|
| all episodes | 100 | metadata universe |
| positive terminal reward | 81 | claimed successes |
| zero terminal reward | 19 | failures without task text |
| task text + positive reward | 46 | eligible feasibility pool |
| task text + zero reward | 0 | no instructed failure comparison |
claim firewall The source label records task success, not a universal no-collateral-motion contract. The fixed complement monitor may audit a declared downstream use policy, but a complement-only alert cannot be called a poisoned source label. Original-label errors, downstream-policy incompatibilities, measurement failures, and abstentions must be reported separately.
adapter gate The existing unmasked complement field would mostly detect the permitted robot. Before scoring, the system needs an authority-fixed robot exclusion, per-view registration and visibility tests, and deterministic multi-view selection or fusion. No label-quality or accuracy number is extracted from this qualitative screen.
one-sided semantics Verified unexpected change outside the actor mask may raise an audit alert. The converse is harder: masking the robot does not prove that its occluded sweep stayed unchanged. No alert becomes PASS only with measured temporal or multiview coverage of the declared protected complement; otherwise it is ABSTAIN. This keeps an offline QA screen from quietly claiming runtime assurance.
| packet gate | result | meaning |
|---|---|---|
| development episodes | 12 | simple single-object placement instructions |
| extracted frames | 48/48 unique | two views × two phases |
| source/image hash verification | valid | source drift and frame tampering fail closed |
| fixed-role grounding | 11/12 locally valid | D010 rejected for a broad robot mask |
| complete TARGET + DESTINATION coverage | 0/12 | no episode reaches outcome-measurement admission |
| outcome calls | 0 | missing visibility remains ABSTAIN, not success |
what bound Among the 11 valid responses, BEFORE coverage is TARGET 10/11 and DESTINATION 11/11; AFTER coverage falls to TARGET 3/11 and DESTINATION 9/11. The 25% robot-overlap threshold is not the dominant blocker. The actor is still holding or occluding the task object, or the role leaves the frame. Reported confidence does not cure absent measurement.
next experiment Keep this material development-only and add a small temporal-coverage pilot before any outcome scoring: search for a post-release frame only under independently verified release, role visibility, and robot separation; otherwise model interruption and occlusion explicitly through a temporal window. Multi-view fusion requires calibrated time/geometry validity. Verified unexpected change may ALERT, but incomplete protected-scene coverage must remain ABSTAIN.
alternatives held for their proper jobs ViFailback is stronger for curated temporal failure diagnosis and measurement-loss cases; SAFE's real rollouts are a direct runtime-monitoring comparator; AHA remains a procedural development comparator. None supplies the missing external preservation authority for free.
Decision record: RESULTS_DROID100_GROUNDING_V2.5.md · DROID100_GROUNDING_PROTOCOL_V2.5.md · EXTERNAL_VLA_DATA_SELECTION_V2.5.md · metadata audit: droid-100-metadata-audit.json · sources: DROID paper, DROID-100 metadata, SAFE, ViFailback.
result The semantic IR compiles 24/28 sealed extracted contracts versus 22/28 under v2.2. It recovers exactly K003 and K021, whose only problem was redundant POSITION/ORIENTATION permission for the already system-authorized target. All 24 IR records validate against the frozen schema.
| gate | result | interpretation |
|---|---|---|
| execution-ready | 22/28 → 24/28 | two target-only interface losses recovered |
| non-target authority widening | 0/28 | learned contracts cannot enlarge the write-set |
| newly admitted | K003, K021 | trusted exact target binding only |
| still rejected | K004, K014, K020, K027 | schema error, missing remainder, or illegal enumeration |
not language understanding The compiler does not infer that “white pedestal mug” means target. An external system profile binds the exact entity ID to the target role. Target-only learned MAY clauses are projected into pre-existing system footprints; mixed target/non-target permissions and mismatched roles are rejected.
architectural consequence Interoperability improves through a monotone projection into a system-owned write-set, not through a more forgiving parser. The learned layer can vary its surface form but cannot author away the protected complement.
Scope: retrospective engineering replay only; prior prospective scores remain frozen. Protocol: CONTRACT_INTEROP_PROTOCOL_V2.4.md · result: RESULTS_CONTRACT_INTEROP_V2.4.md.
headline The proposed realistic smooth-object weakness did not materialize. That is a useful narrowing, not an invitation to search for a stranger object. The project keeps the conservative firewall and ordinary baselines; it does not claim that ISR or a special structural field is needed to handle these movements.
| screen | surface | ordinary-flow evidence | meaning |
|---|---|---|---|
| attempt 1 | white hinged case | 193–245 px supported movement | visually plain but structurally easy |
| attempt 2 | green lid | 176–215 corners · 226–285 px movement | rim, tab, embossing, and shading remain sufficient |
| exploratory v2.3.3 | pale plate | 300 corners · 98/116 px movement | strongly trackable under both lighting manipulations |
| exploratory v2.3.3 | blue cup | 289 corners; valid pair measured 0.98 px | object is trackable, but no valid moved challenge was captured |
what is not earned The branch does not establish illumination-invariant PASS. One v2.3.3 blue intensity pair also contains camera stress, and its other blue pair is effectively stationary; neither can be recast as the missing physical-movement counterexample. The exploratory screen is retained only as the stopping audit, not promoted to a sealed result.
what survives A task-authored permission envelope plus an independent deterministic complement monitor remains the durable architecture. Cheap Lab differencing and sparse flow are valid implementations. Invalid camera or missing measurement evidence still yields ABSTAIN. ISR remains outside the admitted path unless a natural deployment condition demonstrates an advantage over those ordinary controls.
decision Stop reducing object texture. Preserve v2.3 attempts 1 and 2 as development evidence, leave the LIGHT cases unresolved, and redirect future work toward naturally occurring deployment failures—occlusion, entry/exit, deformation, registration loss, or temporal interruption—where the current firewall can be challenged without constructing a pathological prop.
Interpretation: the negative result strengthens the surviving claim by removing an unsupported mechanism story. The work is now narrower, cheaper, and harder to dismiss as an engineered ISR demonstration.
capture validity repaired All 18 photographs are unique. Q01/Q02 are quiet; Q03/Q04 light-only changes stay stationary; Q05–Q08 physical moves are detected; and only Q09 crosses the camera gate. Unlike attempt 1, every admission result is the same in both temporal directions.
| pair | forward registration | reverse registration | frozen disposition |
|---|---|---|---|
| Q03 light direction | 25.17 px · 0.99141 | 22.00 px · 1.00853 | locked · stationary |
| Q04 light intensity | 8.95 px · 1.00454 | 1.36 px · 0.99918 | locked · stationary |
| Q08 move + light | 0.57 px · 1.00017 | 0.06 px · 1.00007 | locked · moved |
| Q09 camera | 389.82 px · 0.91633 | 431.53 px · 1.03122 | camera stress both ways |
pivotal challenge absent Q06 and Q08 produce 176–215 reference corners and 74–85 supported clusters. This is not missing evidence, so it cannot distinguish a safe LIGHT resolver from one that would hide a real but untrackable movement. “Smooth” must be defined operationally by tracking evidence, not appearance.
Verdict: apparatus/registration qualification passed; no resolver admitted. Full record: RESULTS_V2.3_QUALIFICATION_ATTEMPT_2.md.
packet and controls All 18 corrected photographs are unique. Both CAL pairs are quiet, the declared CAMERA pair crosses the gate in both temporal directions, and textured/smooth movements are detected. Q08's moved cluster vetoes the frame despite simultaneous unverified lighting evidence.
| pair | forward registration | reverse registration | frozen disposition |
|---|---|---|---|
| Q03 light direction | 62.03 px · 0.97590 | 29.87 px · 0.99079 | direction-dependent gate |
| Q04 light intensity | 60.23 px · 0.97174 | 39.67 px · 1.01820 | direction-dependent gate |
| Q08 move + light | 15.21 px · 0.99461 | 46.96 px · 0.98194 | direction-dependent gate |
| Q09 camera | 338.32 px · 0.97181 | 302.59 px · 1.04029 | camera stress both ways |
pivotal challenge absent Q06 and Q08 smooth movements are supported at 192.65 and 244.61 full-resolution pixels. The hinged case has usable latches, seams, and corners, and the nearby inch ruler indicates the nominal 1 cm manipulation is closer to inch scale. Thus the capture does not test whether a LIGHT resolver would conceal an unverified physical movement.
Verdict: no candidate admitted. Full record: RESULTS_V2.3_QUALIFICATION_ATTEMPT_1.md.
replicated at one centimetre The direct judge false-passed all eight textured/smooth 1 cm judgments and one forbidden pink-bowl judgment, with no abstentions and 0.95 confidence throughout. Its explanations explicitly asserted that the visibly moved objects remained in place. This extends the natural failure beyond the earlier 3 cm selective case.
| prospective v2.2 arm | decided | correct among decided | false pass | false fail | abstain |
|---|---|---|---|---|---|
| direct Gemini | 20 / 20 | 11 / 20 | 9 | 0 | 0 |
| extracted contract + v2.2 | 12 / 20 | 12 / 12 | 0 | 0 | 8 |
| gold contract + v2.2 | 16 / 20 | 16 / 16 | 0 | 0 | 4 |
| target-only control | 20 / 20 | 10 / 20 | 10 | 0 | 0 |
| unconditional complement | 16 / 20 | 14 / 16 | 0 | 2 | 4 |
| frozen Lab-only ablation | 20 / 20 | 16 / 20 | 0 | 4 | 0 |
authority boundary learned something too BEFORE-only extraction was 25/28 schema-valid and 22/28 execution-ready. Two P contracts added redundant MAY clauses for an entity named WHITE_PEDESTAL_MUG; strict authority rejected them because that alias was not deterministically bound to the system target. This is safe but costly. Target-role binding is a v2.3 interface hypothesis, not a v2.2 rescore.
Verdict: literal safety gates pass, but the scientific specificity claim receives only partial confirmation. V2.2 demonstrates “false failure → abstention,” not illumination invariance. Full record: RESULTS_SPEC_EXECUTION_V2.2.md.
integrity and localization All sealed hashes, all 28 opaque-to-source image copies, and all 14 registrations reproduce. Each textured/smooth violation has one compact proposal over the moved object with supported motion of 200.40–384.89 full-resolution pixels; CAL and target-only SUCCESS produce none. The nine direct false passes and prospective denominators are unchanged.
| post-hoc candidate | decided | correct | false pass | abstain | disposition |
|---|---|---|---|---|---|
| frozen v2.2 | 16 / 20 | 16 / 16 | 0 | 4 | retain |
| stationary track overrides unverified | 20 / 20 | 20 / 20 | 0 | 0 | reject: unsafe aggregation |
| illustrative low-chroma fallback | 20 / 20 | 20 / 20 | 0 | 0 | reject: counterexample |
| unchanged ISR + frozen flow | 17 / 20 | 15 / 17 | 2 | 3 | reject |
Goldilocks trap exposed Both favorable 20/20 replays merely fit the observed LIGHT pairs. The chroma fallback places the real N14 smooth-object movement (0.279) on the same side of its proposed 0.35 boundary as every lighting residual; flow happens to rescue N14 in this capture. A texture-poor unverified movement could therefore be cleared. The permissive aggregation is worse: any stationary background patch could outvote missing evidence elsewhere.
next gate, not next threshold V2.3 needs a system-owned target-role token and a separately defined illumination-validity measurement tested against a deliberately texture-poor physical move that defeats flow, plus held-out lighting types. Until that evidence exists, LIGHT remains ABSTAIN.
Full record: ADVERSARIAL_CHECKPOINT_V2.2_POSTRESULT.md. This is explicitly post-result analysis and does not rescore v2.2.
compiler boundary The exact POSITION+ORIENTATION remainder pair is now normalized to the executor's conservative visual-change proxy with source-clause provenance. Canonical clauses pass through; partial, duplicate, mixed, enumerated, ambiguous, and invalid forms still reject. Replayed acceptance rises from 11/28 to 25/28: P becomes 14/14, S remains 11/14. After blind extraction, an independent authority profile requires the extracted permission set to equal the system-authorized set; extra or missing permissions reject.
measurement boundary Ordinary Lab remains the sensitive proposal field; forward/backward optical flow corroborates physical displacement inside each unexpanded proposal box. Any supported moved cluster fails—even when a larger stationary background cluster is present—any unverified proposal forces ABSTAIN, and only verified-stationary proposals are cleared. The recurring embossed-mug residual measures 1.92–2.20 full-resolution px; true v2.1 collateral proposals measure 143–336 px.
| development replay | decided | correct | false pass | false fail | camera abstain |
|---|---|---|---|---|---|
| v2.1 gold contracts | 16 / 16 | 16 / 16 | 0 | 0 | 4 / 4 |
| v2.1 extracted contracts | 15 / 16 | 15 / 15 | 0 | 0 | 4 / 4 |
| v1.6 unlisted cross-replay | 9 / 9 | 9 / 9 | 0 | 0 | 3 / 3 |
adversarial correction before capture The first freeze was not treated as sacred merely because replay was perfect. The audit found that a dominant stationary cluster could hide a smaller moved cluster, the model still had authority to invent a schema-valid permission, malformed nested values could escape the standalone normalizer, and the reported 30–120 px stability range described v2.1 alone. All four issues were corrected, tested, disclosed, and re-frozen before any v2.2 image existed.
sensitivity, not victory The cross-set threshold sweep is 25/25 correct through 90 px; v1.6 loses one true movement at 120 px, and v2.1 loses two at 150 px. The largest independent no-move flow is 6.65 px. A moved proposal's 89.20 px cluster is not itself an object-level boundary because another supported moved cluster can veto, so no universal safe interval is claimed; 60 px is retained as a development-supported operational policy. The corrected any-moved-cluster rule restores 16/16 at 0%, 15%, and 50% box expansion, but 100% expansion still misses one movement; zero remains frozen. All 80 independent same-state calibration ROIs stayed stationary. A future texture-poor motion may still be unverified; it must ABSTAIN rather than pass.
Prospective test completed: two new layouts added a locked-camera lighting/highlight nuisance, textured and smooth 1 cm moves, the paired selective permission, CAL, SUCCESS, and CAMERA. V2.1 remains unchanged. Repair: DEVELOPMENT_REPAIR_V2.2.md · red-team record: ADVERSARIAL_CHECKPOINT_V2.2.md · protocol: PROSPECTIVE_REPAIR_PROTOCOL_V2.2.md · result: RESULTS_SPEC_EXECUTION_V2.2.md.
what this is The durable proposition is a visual frame-condition firewall. The commanded object and destination define a permitted write-set; registered measurement audits the complement for unauthorized scene change. Learned grounding may translate task symbols into coordinates, but it does not own the invariant or the final verdict. This is an implementation of established runtime-assurance logic, not a claim to have invented monitors or abstention.
occupied ground VLM success detection and free-form failure reasoning are already active lines in Vision-Language Models as Success Detectors and AHA. Sentinel combines a statistical action-consistency monitor with VLM task-progress reasoning. SAFE learns failure scores from VLA internal features and calibrates them with conformal prediction; uncertainty-aware policy steering calibrates a learned verifier to decide whether to act, clarify, or request intervention. “Models can be confidently wrong; monitor them and defer under uncertainty” is motivation, not this project's novelty.
closest distinction Code-as-Monitor is the closest comparator. GPT-4o generates its subgoals and constraint set, a VLM-derived segmenter selects constraint-related entities and parts, and GPT-4o generates executable monitor code. Common thresholds may come from an external knowledge base, but unseen-task thresholds may also come from the VLM. Branch-coverage testing can establish executability while leaving specification completeness unresolved: a valid program may monitor the wrong or incomplete contract. The stricter hypothesis here is that a system-authored permission complement prevents that omission channel.
failure-family correction Sentinel's canonical confident wrong-placement case violates the positive task goal. The photographed hard case satisfies the positive target goal while violating a remainder-of-scene invariant. It is therefore better described as a frame-condition failure, not simply another task-progression failure. Conversely, the four 0.95 Gemini misses used explicitly named protected objects; v1.6 proves recognition-light unlisted detection but did not run direct Gemini as an outcome judge. The strongest conjunction remains to be tested prospectively.
different abstention trigger Selective-prediction systems abstain because a learned answer or policy is uncertain; see also the VQA-focused dual-assessment reliability study. Here the monitor abstains when its measurement preconditions are invalid, even if the learned model is confident. This is established metrological/runtime-assurance discipline—see NASA's formal runtime-assurance framework—but its explicit application to a VLM-authored success claim is the useful design stance.
parallax claim ceiling The empirical result is specific: one homography-compensated image-plane displacement tolerance does not safely span locked-camera and viewpoint-changed scenes with depth variation. It does not establish that Code-as-Monitor's fixed multiview RGB-D pipeline fails. A direct comparison would require equivalent RGB-D trajectories, tracking, and camera conditions; an existing-image pilot must be labeled “CaM-style model-authored monitor,” not a reproduction of CaM.
| track | question | next gate |
|---|---|---|
| A · contract authority | Does a fixed permission complement catch collateral violations omitted by a model-authored monitor? | Seal authored DSL specifications before AFTER images; compare omission and execution errors separately. |
| B · validity | Does a nuisance/observation gate turn unsafe geometric decisions into calibrated abstentions? | Manipulate camera/extrinsic, visibility, depth, and lighting conditions without sacrificing locked-camera coverage. |
| C · ISR invariance | Can a sparse structural channel resolve dense photometric alarms cheaply when invoked locally? | New candidate, calibration, and outcomes; no retuning of the rejected v1.7 field. |
| D · ISR relations | Do temporal islands or corridors expose consequential relation/clearance changes beyond ordinary scene graphs? | Choose a genuinely relational task and beat strong tracking, segmentation, and collision-geometry controls. |
pilot result Gemini's 16/16 BEFORE-only specifications selected the target goal, full permission complement, safe missing-evidence behavior, and calibration rather than invented thresholds. Yet its separately prompted direct judge false-passed all 8/8 collateral violations at 0.95, false-failed both camera-stress controls at 0.95, and never abstained. The simple contract-omission hypothesis is weakened; the sharper observed failure is specification–execution inconsistency: the learned workflow represented the invariant but its free-form judge did not enforce it.
adversarial qualifier Primitive selection is not complete monitor quality. Six specifications unnecessarily permit a transit sweep, M12 binds destination occupancy anomalously, and M05 adds an unstated orientation constraint. Nine of sixteen retain the minimal source/destination permission pattern without the transit expansion or anomalous binding. These are disclosed post-hoc diagnostics, not changes to the preregistered primary classification.
compiler result A strict two-state compiler makes those semantic defects operational. Literal B1 is ready in 0/16 because every value correctly awaits calibration. With the frozen structure-only policies, only 9/16 contracts are valid and 8/16 execution-ready: six transit permissions are rejected as unobservable, one destination binding is rejected, and one otherwise-valid contract requests an unsupported orientation primitive. The compiler abstains or rejects; it never silently deletes a clause.
v2.1 sealed prospective result Twenty-eight uniquely hashed photographs produced 28 BEFORE-only extractions, 20 direct non-CAL judgments, and four measured arms. The gold firewall scored 13/16 task outcomes with zero false passes; the direct judge scored 7/16 with nine false passes. Target-only produced ten false passes, while an unconditional complement produced four false fails and no correct selective label flip. The gold selective monitor achieved the required P→FAIL/S→PASS flip in one of two layouts.
v2.1 falsifier retained The gold executor false-failed the Set A nominal success under both contracts and its selective pink case under S. Those three errors are attributable to the ordinary change surface, not extraction. They prevent the claim that the compiled arm already dominates without an offsetting cost. Seventeen extracted-arm rows also terminate at compilation, so its zero false passes come with unusably low coverage.
| v2.1 arm | task result | unsafe false pass | false fail | camera abstention |
|---|---|---|---|---|
| direct learned judge | 7 / 16 | 9 | 0 | 0 / 4 |
| extracted + compiled | 6 / 7 decided | 0 | 1 | 4 / 4 |
| gold contract + executor | 13 / 16 | 0 | 3 | 4 / 4 |
| target-only control | 6 / 16 | 10 | 0 | 4 / 4 |
| unconditional complement | 12 / 16 | 0 | 4 | 4 / 4 |
Sequencing: Track A has prospective support for the architecture and a newly exposed compiler-normalization problem. Track B now has a concrete residual: remove Set A incidental-change false fails without losing OTHER detection, CAL quietness, or camera abstention. ISR remains separately gated and may not enter as a post-hoc v2.1 repair. V2.1: protocol · result.
narrow prospective advantage Thirty-two new photographs form 16 distinct pairs across plain-sheet and repeated-wood layouts. Hashes and the v1.7 thresholds were frozen before outcome scoring. Gemini saw only the target mug and target region; collateral identities and boxes were schema-forbidden.
| v1.6 method | raw collateral alarms | CAL false alarms | median moved-sweep coverage | median runtime |
|---|---|---|---|---|
| ordinary Lab | 8 / 8 | 0 / 4 | 70.2% | 39.97 ms |
| full-frame SIFT | 7 / 8 | 0 / 4 | not scored | 165.59 ms |
| ISR mass residual | 8 / 8 | 0 / 4 | 9.0% | 177.60 ms |
Registration found the two intended camera shifts and one accidental shift in A_PLAIN_1CM; all three abstain before verdict. Both methods decide the remaining seven collateral failures correctly, including the texture-poor 1 cm case in layout B. ISR passes both nominal successes. Ordinary Lab false-fails B_SUCCESS because a 6,176-pixel residual forms on a stationary embossed white mug.
The favorable error was challenged post hoc. The residual does not intersect the target envelope, and global registration is clean. Lucas–Kanade flow directly inside it finds 41 coherent points with only 1.78 full-resolution px displacement; expanded-crop SIFT and flow estimates remain below 1.87 px, versus the frozen 60 px object-motion tolerance. The evidence supports photometric/texture sensitivity rather than hidden mug movement.
protocol correction The sealed runner omitted the protocol's frozen v1.5 full-frame SIFT arm. Run unchanged after discovery, SIFT has 0/4 CAL alarms, detects 7/8 raw and 6/7 camera-admitted collateral moves, and false-passes A_PLAIN_3CM. The CAMERA captures also omit their prescribed target move; they remain camera controls but are excluded from task-outcome denominators.
Claude's blind return validated all 16 neutral pairs and supplied 32 inch-ruler intervals (median 96.35 px/cm at the ruler plane), but also claimed five low-confidence shifts of the textured mug from unregistered differences. Registered SIFT and flow disprove all five at 0.12–2.87 px versus the 60 px movement tolerance. This includes N03/B_SUCCESS, so the ordinary error remains a genuine photometric false alarm. Claude also missed the accidental A_PLAIN_1CM camera stress, reinforcing that a second model is annotation evidence, not physical ground truth.
On seven camera-admitted collateral pairs, both dense methods overlap the verified moved-object polygon 7/7. Lab has 70.2% median swept-object coverage and 0.385 IoU; ISR has 9.0% coverage and 0.086 IoU, though its sparse detections have higher precision. Median local runtime is 39.97 ms/pair for Lab, 165.59 ms for SIFT, and 177.60 ms for ISR.
Dialectical result: the ISR field contains a real photometric-invariance clue, but it wins none of the declared admission routes: no movement missed by all ordinary arms, worse localization coverage, equal abstention coverage, and 4.44× Lab cost. It also falls from 8/8 to 5/8 collateral detections at a +10% score threshold while Lab remains 8/8 through +20%. The current candidate is closed without tuning away either its sparse geometry or ordinary's one false alarm.
Surviving direction: target-conditioned recognition-light scene integrity is viable. For deployment, ordinary registered change is the stronger starting control; islands, corridors, or cheaper structural invariance would be separately preregistered hypotheses, not repairs to v1.7.
recognition-light miss The full-frame control receives no protected-object identities or boxes. It searches outside only the task-permitted source/destination envelope for SIFT matches that cluster jointly in position and residual-motion vector.
The missed bowl provides five coherent matches. Lowering the cutoff to five would also admit no-move repeated-texture clusters, so the miss cannot honestly be tuned away. Object-scoped SIFT and flow both detect it because semantic ROIs suppress the background ambiguity.
What this supports: recognition-light discovery is genuinely harder than verifying enumerated objects. What it does not support: that ISR islands are the solution. Dense motion proposals, better background suppression, segmentation, and standard multi-object tracking remain competitors.
Next gate: a newly captured, genuinely unlisted collateral object with frozen false-alarm calibration and identical ordinary/ISR abstention rules. Existing-data box withholding is only a simulation.
second ordinary solution Forward/backward pyramidal Lucas–Kanade flow tracks protected-object corners after whole-frame homography compensation. All 80 locked calibration tracks verify; the worst stationary residual is 6.65 full-resolution px.
Flow reproduces SIFT with zero false passes or false fails and ten abstentions. Registered local change adds one wrong-object decision but false-fails photo-v0.2 IMG_2 and photo-v0.3 IMG_4: both contain small bowl shifts above the no-move noise floor but below the 60 px operational tolerance.
Dialectical result: sensitive change is not automatically useful compliance. A twice-noise threshold answers “did anything measurably change?” while the task requires a declared answer to “how much change counts as disturbance?” Changed fraction lacks physical units, so it cannot settle that policy alone.
ISR gate: SIFT and flow both solve enumerated-object preservation. The next experiment must withhold a semantic box from a collateral object and compare full-frame ordinary motion clustering with ISR islands. Repeating the bowl test would add volume, not information.
ordinary method sufficient A SIFT ROI tracker uses cached semantic reference boxes, stored camera homographies, and coherent residual-vector clusters. Weak tracks are unverified rather than silently stationary. Cached Gemini boxes still supply target placement.
The no-move feature-tracking maximum is 6.68 full-resolution px, but sensor noise is not the task tolerance. A twice-noise cutoff falsely rejects two nominal successes containing small bowl residuals of about 34 and 19 px. The independently derived 60 px candidate separates these incidental shifts from deliberate collateral moves without approaching the observed decision boundary.
Implication: the repeated collateral-bowl false pass does not justify ISR. Once identity is initialized, standard local correspondence catches all four instances. ISR's remaining fair test is recognition-light unlisted change, fragmented or texture-poor regions, or efficient attention under parallax—against flow, segmentation, and modern tracking.
Scope: retrospective algorithm design, correlated tabletop images, semantic initialization and placement from Gemini, camera frames abstained, and no open-world object. This is a baseline result, not a robotics novelty claim.
material bug corrected The Claude audit applied a 900 px homography using longest-dimension scale even though registration resizes to fixed width. EXIF-rotated portrait images made the mismatch material. Corrected locked residuals are 2.35 px median, 9.17 px p95, and 28.02 px maximum; corrected camera residuals are 28.03 px median, 81.05 px p95, and 214.85 px maximum.
Not Goldilocked on preservation: every locked photo-v0.4/v0.5 verdict is unchanged for any tolerance from 8.09 px inclusive to 149.13 px exclusive. The corrected independent-model maximum is 28.02 px, leaving a broad observed interval. A 60 px candidate comes only from doubling that maximum and rounding; future tolerance must be tied to a declared minimum motion and loss function.
Oracle removed: camera abstention previously trusted the manifest's expected camera flag. It now uses measured registration, and a regression test proves that deliberately false manifest flags cannot suppress abstention. The current outputs do not change because measured and declared conditions agree.
Claims narrowed: placement is only target-center-inside-full-mug-box, not base-centering metrology. Every protected object is explicitly enumerated, so open-world collateral change is untested. Camera calibration strongly separates large perturbations but sparsely samples the mild 34–50 px boundary. V0.4 IMG_4's camera label followed the measurement and is not independent detector validation.
Buried direction: the current verifier solves closed-set preservation. The still-open, ISR-relevant problem is recognition-light unexpected change for regions that were never named or grounded. Ordinary flow, ROI tracking, and segmentation must define that residual first.
independent disagreement Claude Opus 4.8 received a sealed folder with 32 original images, neutral labels, schema, and validator—but no Gemini boxes, thresholds, or results. It returned 192 validated object records after HSV segmentation, visual verification, and manual corrections.
After the v1.2 scale correction, three locked residuals still exceed 20 px: Blue-A's cup at 22.06 px and Blue-C's bowl at 20.05 and 28.02 px. None exceeds 60 px. Thus the rejection of 20 px survives, while 60 px remains explicitly heuristic.
Median cross-model agreement is tight, but six large pedestal-mug discrepancies arise because the arithmetic box center changes depending on whether the handle and full base count toward the visible extent. The correction is conceptual as well as numerical: a box center is an annotation convention, not a unique physical object center. Future metric work should declare a contact point or object-specific anchor.
Protocol correction: the visible scale is labeled in inches. Claude correctly returned all original ruler fields as ambiguous because the packet demanded centimeter ticks. V0.9 stays unchanged as executed; it is not retroactively rescored.
Interrupted addendum: before its token limit, Claude saved 3-inch/8-inch points for all 32 frames but not the requested final files. The 24 reference/locked frames yield 219.64 px/in (86.47 px/cm) median. This contextualizes 28.02 px as roughly 0.32 cm and 60 px as roughly 0.69 cm, but does not create a global scale for raised objects across the perspective field.
architecture survives The user confirmed that photo-v0.4 and photo-v0.5 followed the v0.3 file order. Roles and the operational rule—20 px under locked capture, ABSTAIN on detected camera stress—were written before either set was inspected or sent to Gemini.
Both direct errors reproduced the exact hard mechanism: correct mug placement plus collateral pink-bowl motion. Direct Gemini said the protected objects appeared unchanged; the grounded path kept placement true but measured registered bowl displacement of 149.1 px and 189.3 px and returned FAIL. Across photo-v0.2 through v0.5, direct Gemini is now 24/28, with all four errors on this one mechanism.
Photo-v0.4 IMG_4 unexpectedly crossed the pre-existing camera gate (scale 0.977; full-resolution translation 50.2 px). Its frozen task-success label was retained, but the grounded verdict became ABSTAIN. This is a nuisance-condition correction from declared measurements, not outcome relabeling.
Scope: these are correlated related-set confirmations, not fourteen independent trials. The 20 px rule is preserved as the prospective execution condition; the later blind audit rejects it as a future cross-annotator threshold. Grounded correctness is conditional on 64.3% coverage.
identity stable All 32 calibration frames returned all six requested boxes: 192 boxes, no malformed outputs, no abstentions. The similar white mugs remain distinct, and the blue cup does not collapse into the blue target under locked capture.
Blue-D exceeds the old tolerance on a camera-only frame. Overlay audit confirms correct mug, cup, target, bowl, and apple identities. Multiple raised objects shift differently after whole-frame registration: this is depth-dependent parallax, not a semantic swap or a moved object.
Design correction: do not inflate one global threshold. For locked capture, carry a provisional 20 px tolerance—twice the observed maximum rounded upward. For detected camera stress, return ABSTAIN on preservation until per-object tracking with uncertainty, depth-aware reprojection, stereo, or another measurement surface is available.
Protocol deviation: grounding was run before independent annotations were sealed. Cached boxes must now be withheld from a separate human annotator; annotation by the same agent that audited them would not be blind.
At this stage: 20 px remained a candidate until independent boxes and ruler endpoints could quantify absolute error and physical scale. No photo-v0.4/v0.5 outcomes had been uploaded or read into this calibration; it was subsequently frozen as the provisional operational rule for v0.9.
both sets usable The White and Blue rearrange folders each contain four materially different layouts with an empty target, reference, two no-object-move repeats, and one camera-only perturbation. The similar white mug replaces the earlier black distractor. The Blue condition replaces the white case with a blue cup, deliberately testing target-color versus object-identity grounding.
All 16 nominally locked repeats remain below the established camera-stress gate; all eight frame-4 perturbations exceed it. Every comparison remains registerable. Small natural camera/exposure variation in frames 2/3 is retained as realistic calibration noise rather than sanitized away.
Admission decision: use both sets. The blue cup is a useful hard condition, not a protocol violation. At this stage photo-v0.4 and photo-v0.5 were inventoried as two additional eight-frame outcome sets and deliberately left unevaluated until the role and camera-policy rule was explicit; v0.9 records their later evaluation.
Still pending: independent human boxes/centers and ruler endpoints. Grounding is now complete, but registration plus learned boxes cannot independently prove that every object was physically motionless; immobility currently rests on the capture protocol plus visual review.
threshold survives The v0.2 wrong-object repeat and v0.3 success repeat were treated as same-state localization-noise controls. Each first frame was registered to its repeat; clean-prompt Gemini centers were then compared through the homography.
The worst repeat drift doubled is 64.38 px, just below the frozen 70 px tolerance. The smallest actual collateral drift is 156.25 px. No threshold or label was refit. Nearly all non-bowl repeat residuals are below 4 px, implicating irregular-object box variability more than global registration.
Limit: only two same-state pairs, one capture regime, and model-produced boxes. Pixel tolerances change with range, focal length, crop, and resolution. The next real calibration requires true no-move repeats and independent annotations gathered before outcome scoring.
exact mechanism repeat A second eight-photo capture replaced the patterned distractor mug with a small black mug, enlarged the target, reshot the layout, and added a third passing outcome. The 70 px preservation tolerance and model settings were unchanged.
On v0.3 IMG_3, direct Gemini again returned PASS at 0.95 confidence and said the pink bowl appeared unchanged. Registered grounding measured the bowl 187.3 px from its expected center and failed preservation while passing target placement. The three v0.3 successes topped out at 28.6 px protected-object residual.
Cross-version sensitivity: under the cleaned grounding prompt, the largest success residual is 39.4 px; the smallest true collateral residual is 156.2 px. Every threshold strictly inside that interval yields the same thirteen decided grounded labels. No threshold was adjusted for v0.3.
audit correction The initial grounding prompt carried an irrelevant synthetic-era sentence about “large and small blue squares.” After removing it and versioning the cache, both photo sets were regrounded. v0.2 remained 7/7; v0.3 IMG_6 returned an impossible pink-bowl box and correctly triggered ABSTAIN. The frame was not retried. The earlier 14/14 observation is preserved as prompt-contaminated history, not the standing result.
Limit: this is a stronger mechanism replication, not independent benchmark accuracy. The room, table, task grammar, target object, and several protected objects are shared; labels were inspected; boxes remain model-produced rather than human annotated.
mechanism replicated Eight new iPhone photographs formed one common BEFORE frame and seven correlated AFTER outcomes. Roles were frozen before the API call: two successes, a target success with collateral bowl movement, two wrong-object repeats, a wrong object plus collateral movement, and success/failure camera-shift variants.
| frame | mechanism | truth | pixel delta | direct | grounded |
|---|---|---|---|---|---|
| IMG_2 | success | pass | fail | pass | pass |
| IMG_3 | success + pink bowl moves | fail | fail | pass | fail |
| IMG_4 / IMG_5 | wrong floral mug · repeat | fail | fail | fail | fail |
| IMG_8 | wrong mug + bowl moves | fail | fail | fail | fail |
| IMG_7 | success + camera shift | pass | fail | pass | pass |
| IMG_6 | wrong mug + camera shift | fail | fail | fail | fail |
The direct miss was explicit, not ambiguous: on IMG_3 the model said the pink bowl appeared unchanged and returned PASS at 0.95 confidence. Under the cleaned grounding prompt, the grounded path measured a 207.9 px registered bowl displacement and failed the preservation predicate. Positive frames had worst protected-object residuals of 37.4 and 39.4 px; collateral frames were 156.2 and 207.9 px. The provisional 70 px threshold therefore sits in a broad observed gap, although it is not yet independently calibrated.
Limit: one inspected layout, correlated frames, visually distinct mugs, developmental margins, and learned boxes rather than human annotations. This is mechanism evidence, not a 7-trial accuracy estimate. The negative pixel-delta result matters: parallax from small viewpoint changes creates residuals around stationary raised objects, so whole-frame change is an attention cue, not yet an object-compliance measure.
claim narrowed A full code-and-artifact audit found that the conceptual direction survived but the v0.2 harness overstated nuisance coverage and experimental independence. Historical artifacts remain untouched; the corrected generator and local artifact are versioned v0.3.
Corrected: exact static background replay; clutter in both frames; unique controls; distinct corridor failures; RGB-detected target map instead of oracle coordinates; raw-response-only cache with current predicate reevaluation; strict grounding validation; signed residual convention; measured rather than ground-truth legacy fusion; media hashes and stronger manifest checks.
Still open: no new Gemini calls validate the corrected generator; the original nuisance gate remains unmet; flat decoy is an explicit RGB-equivalent information-limit construct rather than a photorealistic print; box-level depth attribution, real calibration, prompt robustness, and held-out multi-object rollout data remain unresolved.
new Existing Robotics-ER photographs were not discarded or casually folded into a new claim. They were normalized into a 126-record manifest with explicit provenance, task endpoint, media, ground truth, and limitations. All are marked legacy exploratory · not held out.
| reported result | machine-readable selection | correct |
|---|---|---|
| parallax overall | all surface_height records | 99 / 113 |
| clean non-empty | non-empty · no detect failure · finite parallax | 93 / 103 |
| stored direct VLM | prior model result exists | 17 / 27 |
| retrospective appearance + surface | prior result · clean detection · raised/flat | 24 / 24 |
| stored placement agreement | placement + prior verdict | 13 / 13* |
* Only one positive placement case. This is not balanced placement accuracy. The 24/24 fusion excludes three detector failures, is retrospective on development data, and does not solve a printed semantic decoy raised on a real box.
Historical record: the initial synthetic experiment removed printed IDs, added similar distractors, metric language, matched latent initial states, and classical baselines. The API slice itself used only clean index-000 scenes; nuisance diversity was present in the local corpus, not demonstrated for Gemini.
| condition | coverage | decided accuracy | false pass |
|---|---|---|---|
| oracle geometry | 27 / 27 | 100.0% | 0 |
| classical RGB components | 24 / 27 | 100.0% | 0 |
| classical components + depth | 27 / 27 | 100.0% | 0 |
| Gemini direct | 27 / 27 | 92.6% | 2 |
| Gemini boxes + RGB predicates* | 24 / 27 | 100.0% | 0 |
| Gemini boxes + predicates + depth* | 27 / 27 | 100.0% | 0 |
* Historical grounded conditions used an oracle target rectangle for containment. Corrected v0.3 detects the target from RGB.
Use a learned model to tell us what might matter and where it is; use independent, declared measurement surfaces to decide whether the task's physical and metric predicates are actually satisfied; return ABSTAIN whenever the required surface is absent.
The potential spatial-field contribution is not another left-of function. It is a recognition-light account of what else changed, what relates to what, and where expensive verification should look next—but only if it outperforms ordinary baselines on real rollouts.
Robotics-ER had been run as a mostly single-object, height-above-plane instrument. ISR and later VTL work had moved toward islands and relational scene structure because whole-frame mass becomes confounded when several objects or clutter share the image. The question was not whether to merge the repositories. It was whether the broader spatial work had a credible robotics home.
The risk was obvious: robotics already has excellent deterministic geometry, tracking, occupancy, collision checking, SLAM, and scene graphs. An outsider can easily reinvent a weaker version and mistake unfamiliarity for originality. So this project was deliberately split out and built as a sequence of gates.
A small independent package keeps Robotics-ER and ISR untouched while their ideas are tested against ordinary controls.
Containment, metric left-of, clearance, unchanged-object, height-above-plane, clear corridor, and preserve-scene-except-target. Missing evidence never defaults to success.
Six families; clean and cluttered successes; wrong relation; wrong similar object; flat decoy; named/unnamed collateral movement; corridor obstruction.
Renderer-specific by design and explicitly not a real-world segmenter. Their job is to prevent trivial predicate arithmetic from being misidentified as a contribution.
Gemini either decides end-to-end or returns object boxes; identical deterministic predicates consume the boxes, with and without independent height.
Content-addressed raw-response cache, current-code reevaluation, pass/fail family-balanced smoke slices, targeted scene IDs, repeat stability, predicate residuals, and false-pass metrics.
126 real records normalized without pretending they are new or held out. Media hashes and validation protect provenance; a separate audit writes every denominator's exact trial IDs.
One five-object layout tests success, wrong object, collateral disturbance, repeats, and camera shift through pixel delta, direct judgment, and grounded deterministic relations.
Two reserved related sets replicate the collateral-motion miss; locked-camera grounded decisions remain correct while parallax-triggering frames abstain.
| family | claim being audited | hard control |
|---|---|---|
| place_inside | large blue square inside target and raised | flat decoy · similar small-blue object |
| left_of | 3.5 cm edge-to-edge margin | wrong object · near metric miss |
| separate | 9 cm minimum clearance | visually plausible insufficient gap |
| move_without_disturbing | target move while named B remains fixed | named collateral movement |
| move_preserve_scene | target move while every known tracked object remains fixed | unnamed C moves |
| clear_corridor | straight path clear of every blocker by 1.8 cm | intervening island/object |
| candidate claim | current reading | why |
|---|---|---|
| New relation geometry | do not claim | Classical components + existing geometry solve the synthetic controls. |
| Learned judge is enough | rejected | Stable false passes on height and metric/identity controls; real flat-print errors already exist. |
| Learned perception is useless | rejected | Gemini boxes grounded objects well; the error lived in unconstrained verdict synthesis. |
| Independent measurement + abstention | supported narrowly | No wrong decisions in the historical grounded conditions; the first real multi-object grounded run was 7/7 while direct judgment false-passed collateral movement. |
| ISR mass-field delta improves motion auditing | reopened narrowly | It lost 0/4 versus 4/4 retrospectively, then matched 8/8 prospective collateral alarms, matched 0/4 CAL false alarms, and avoided one ordinary photometric false failure. Localization and replication are still absent. |
Candidate A · islands as recognition-light multi-track. Segment the field into stable regions, then attach ordinary geometry or parallax per island. Anchor/satellite roles may provide task-conditioned priority without requiring semantic recognition everywhere.
Candidate B · corridors as relational interference. A corridor between islands can express “between” and obstruction, but must be compared with standard occupancy/collision geometry.
Candidate C · layered field delta as attention gate. The first retrospective test was negative: edge/tone/chroma residuals fragmented moved objects while ordinary Lab detected 4/4. The frozen prospective capture changes that reading: both methods alarm on 8/8 declared collateral moves, while ISR alone ignores one stationary object's appearance shift.
gate failed The mass-residual + connected-area candidate does not enter the robotics path. It has one prospective specificity advantage, but no movement missed by all ordinary arms, 9.0% versus 70.2% moved-object coverage, 0.086 versus 0.385 IoU, equal abstention coverage, 4.44× Lab runtime, and marked +10% threshold sensitivity. Photometric invariance may be studied separately, but cannot rescue this candidate under the frozen gate.
| decision | status | reason |
|---|---|---|
| Merge Robotics-ER and ISR now | rejected | Would entangle hypotheses before contribution is established. |
| Start with ISR islands/corridors | rejected | Ordinary baselines must define the residual first. |
| Treat whole-frame mass as object location | rejected | Confounded by clutter and multi-object scenes. |
| Print object IDs in the benchmark | removed | Turned semantic grounding into OCR/token matching. |
| Let RGB infer physical height from shadow | rejected | The flat-decoy false pass demonstrates why appearance is not independent height evidence. |
| Count abstention as an embarrassing miss | rejected | Abstention is the correct result when the required measurement surface is absent. |
| Use legacy real data as held out | rejected | It informed prior method development; provenance is explicit. |
| Keep a third independent project | chosen | Allows clean ablation, ordinary controls, and easy abandonment of invalid ideas. |
| Promote ISR mass-field delta | rejected | It matches 8/8 at the frozen point and avoids one photometric error, but localizes far less, costs 4.44× Lab, and drops to 5/8 at +10% score threshold. No admission route survives. |
Held for a separate cleanup conversation. They matter if code or claims are lifted into this project, but were not silently changed in Robotics-ER.
| bookmark | implication here |
|---|---|
CLI defaults to isolate=none | Never make whole-frame measurement the implicit real-scene path. |
| μ terminology drift | Do not inherit a name whose paper and implementation measure different constructs. |
| 8D / 9D kernel naming drift | Freeze dimensional terminology before citing or lifting a kernel. |
| 86/89 selection was prose-only | The new 93/103 rule is executable; it does not reconstruct or validate the historical 86/89 subset. |
| Coverage gaps in CLI/real pipeline/calibration/malformed/stereo | The new runner tests cache, selection, controls, abstention, and manifests; real calibration and stereo regression remain open. |
V2.1 is sealed complete. V2.2 is frozen before new capture and tests whether the two post-hoc repairs generalize rather than merely replay their development sets.
| what | command |
|---|---|
| tests | python3 -m unittest discover -s tests |
| local 112-scene audit | python3 -m relational_auditor.run --per-case 4 |
| import legacy real material | python3 -m relational_auditor.real_manifest --import-legacy-root "/path/to/Robotics-ER 1.6" --manifest data/real-v0.3/legacy-manifest.jsonl |
| legacy selection audit | python3 -m relational_auditor.legacy_audit |
| registered photo baseline | python3 -m relational_auditor.photo_dry_run |
| cached photo Gemini audit | python3 -m relational_auditor.photo_gemini --env-file /path/to/.env |
| same-state repeat calibration audit | python3 -m relational_auditor.photo_calibration |
| independent calibration grounding | python3 -m relational_auditor.calibration_grounding --env-file /path/to/.env |
| blind annotation audit | python3 -m relational_auditor.annotation_audit |
| preliminary ruler scale | python3 -m relational_auditor.ruler_scale_audit |
| adversarial checkpoint | python3 -m relational_auditor.adversarial_checkpoint |
| conventional ROI tracker | python3 -m relational_auditor.conventional_tracking |
| ordinary flow/change controls | python3 -m relational_auditor.ordinary_motion |
| full-frame motion clustering | python3 -m relational_auditor.full_frame_motion |
| unlisted capture template | python3 -m relational_auditor.unlisted_protocol |
| ordinary vs ISR-derived change fields | python3 -m relational_auditor.isr_change_fields |
| prospective unlisted evaluation | python3 -m relational_auditor.unlisted_evaluate |
| post-hoc favorable-result audit | python3 -m relational_auditor.unlisted_adversarial_audit |
| blind v1.6 annotation adjudication | python3 -m relational_auditor.unlisted_annotation_audit |
| verified-polygon localization | python3 -m relational_auditor.unlisted_localization_audit |
| frozen-arm runtime | python3 -m relational_auditor.unlisted_runtime_audit |
| restored full-frame SIFT arm | python3 -m relational_auditor.unlisted_full_frame_evaluate |
| threshold sensitivity | python3 -m relational_auditor.unlisted_sensitivity_audit |
| authoritative v1.8 checkpoint | python3 -m relational_auditor.unlisted_adversarial_checkpoint |
| validate one model-authored specification | python3 -m relational_auditor.model_authored_monitor /path/to/spec.json |
| preflight frozen authoring packet | python3 -m relational_auditor.model_authored_monitor_author --dry-run |
| sealed BEFORE-only authoring stage | python3 -m relational_auditor.model_authored_monitor_author --env-file /path/to/.env |
| sealed direct outcome arm | python3 -m relational_auditor.model_authored_monitor_outcomes --env-file /path/to/.env |
| strict v2.0 compiler audit | python3 -m relational_auditor.model_authored_monitor_compile_audit |
| validate v2.1 prospective packet | python3 -m relational_auditor.spec_execution_protocol |
| verify v2.1 pre-capture hashes | python3 -m relational_auditor.verify_spec_execution_freeze |
| admit completed v2.1 capture | python3 -m relational_auditor.spec_execution_admission --image-root /path/to/capture |
| verify opaque v2.1 extraction packet | python3 -m relational_auditor.spec_execution_extraction_packet --verify |
| finalize sealed v2.1 arms | python3 -m relational_auditor.spec_execution_finalize |
| replay disclosed v2.2 development repair | python3 -m relational_auditor.spec_execution_v22_development --sensitivity |
| verify v2.2 method freeze | python3 -m relational_auditor.verify_v22_freeze |
| validate v2.2 neutral capture packet | python3 -m relational_auditor.spec_execution_v22_prepare --verify-packet |
| sealed v2.2 BEFORE-only extraction | python3 -m relational_auditor.spec_execution_v22_contract_author |
| sealed v2.2 direct outcomes | python3 -m relational_auditor.spec_execution_v22_direct |
| sealed v2.2 measurement arms | python3 -m relational_auditor.spec_execution_v22_measure |
| finalize v2.2 denominators | python3 -m relational_auditor.spec_execution_v22_finalize |
| replay post-result v2.2 adversarial checkpoint | python3 -m relational_auditor.spec_execution_v22_adversarial |
| audit v2.3 qualification capture | python3 -m relational_auditor.v23_qualification_audit |
| audit v2.3.2 qualification capture | python3 -m relational_auditor.v23_qualification_audit --image-root images_rp/v2.3.2 --output-root artifacts/v2.3.2-qualification |
| audit contract interoperability v2.4 | python3 -m relational_auditor.contract_interop_v24 |
| verify frozen DROID-100 packet | python3 -m relational_auditor.droid100_adapter --verify-root artifacts/droid100-feasibility-v2.5 |
| preflight frozen DROID-100 grounding | python3 -m relational_auditor.droid100_grounding --dry-run |
| replay DROID-100 temporal grounding cache | python3 -m relational_auditor.droid100_temporal_grounding |
| replay DROID-100 wrist grounding cache | python3 -m relational_auditor.droid100_wrist_grounding |
| integrate DROID-100 surfaces | python3 -m relational_auditor.droid100_multisurface_integrate |
| rebuild this book | python3 assemble_lab_book.py |