Parallax Metrology
Lab Book · Embodied Verification

VTL Spatial — Robotics Success Verifier

From a silently-confounded demo to a conformant Parallax instrument, with a real frontier-VLM head-to-head.
dated 2026-06-22  ·  repo: Robotics-ER 1.6/  ·  engine: vtl_spatial  ·  deterministic, no learned judge
all work confined to this folder  ·  10/10 tests green  ·  no secret committed (.env git-ignored)
conditions: no physical robot · sandboxed network (Gemini API reachable) · no real rollout frames · key supplied via .env

This is the backward trail: what we started with, what was wrong with it, what we built, what we proved, where we honestly lose, and how the science grounds it. Every number here is reproducible from the scripts named beside it.

STANDING Where this actually is — one snapshot (2026-08-13)

The dated cards below are the chronological trail and can differ in tone; this block is the authoritative "where we are."

proven A deterministic two-view-parallax cross-check discriminates real 3D objects from flat prints on real pixels, across many object types, within a stated envelope. A real frontier VLM (Gemini Robotics-ER) is empirically fooled by flat decoys the instrument catches — 10 cases, some where it names the flatness and passes anyway. The synthetic mechanism is pre-registered, scaled (1200 scenes), with independence + null protocols passing.

scoped The instrument measures height-above-plane and nothing else. Out of envelope: transparent objects, and objects too short to parallax (tape roll, box lying flat). The decoys are constructed adversarial cases — this shows the failure mode exists and geometry catches it, not how often it occurs in the wild.

the edge, precisely A placement-precision test (real cup, offset ladder) came back a clean negative — the VLM localizes to ~0.5cm and judges placement sharply, so the instrument gains nothing on geometry the VLM can see. That agreement is expected (embodied systems localize well) and it is the calibration half of a working detector: a label-quality filter needs both agreement-when-fine (low false alarms — placement) and disagreement-when-wrong (the decoy catch). The instrument disagrees selectively, only when perception is fooled — which is what makes its disagreements trustworthy. The edge is perceptual independence, not out-measuring geometry.

open (1) Real-world value on natural silent failures in real VLA data — the central open question (the collaboration). (2) σ(Ω) temporal axis — needs video. (3) Two arguments still to write for the paper: the learned-judge counterargument (information limit + un-gameability, as two distinct claims) and the falsifiability frame (an information theorem plus a falsifiable monotonic-robustness claim).

2026-08-13 Update — real frames reproduce the mechanism

new The cup shoot happened. Two-view parallax on an ordinary phone (two photos per scene, ~10 cm apart), 7 real scenes, one per class, ground truth read straight from the filenames. The synthetic mechanism reproduces on real pixels: the flat printed decoy reads flat, the real mug reads raised, and the raised decoy (on a box) fools the single depth surface — exactly as predicted.

7 / 7
mechanism reproductions · 1 scene/class, not a statistic
flat
printed decoy correctly caught (−9.2px)
raised
box-decoy trap reproduced (−31px)
n=1
per class — mechanism, not a statistic
flat paper cup (decoy)
9.2
flat green L
11.2
clear glass
19.0
cup on the box
31.3
real mug (right)
45.2
real mug
46.0
|relative parallax| in px · threshold 15.1 · amber = at-plane (flat) · green = above-plane (raised)
Two-frame overlay of a real mug: the orange object marker shifts more than the blue table-plane markers between the two views, and the mug reveals its side wall.
Real mug → RAISED. The object (orange) shifts ~46px more than the table plane (blue) between the two camera positions, and reveals its side wall. That extra shift is real height.
Two-frame overlay of a flat paper cutout of a cup: the object marker tracks the table-plane markers, and the cutout looks identical in both views.
Flat paper cup → FLAT. The object tracks the table plane to within ~9px and looks identical in both frames. A flat thing cannot fake the shift — the whole mechanism, in two pictures.
scenerel.pxverdictexpectedok
bare tableno object
flat paper cup (decoy)−9.2FLATflat
flat green L−11.2FLATflat
clear glass+19.0RAISEDraised
cup on the box (raised decoy)−31.3RAISEDraised
real mug (right)−45.2RAISEDraised
real mug−46.0RAISEDraised

suggestive only The clear glass read raised (+19px, but only 4px over threshold and sign-flipped vs everything else — the shaky one). Tempting to call it "stereo beats active depth on transparency," but that is one anomalous data point, and batch-2 contradicts it: there the glass mostly detect-failed. Honest status — transparency is the instrument's hard case, not a win. Left in only as a hypothesis that would need its own test.

honest Caveats kept in view: n=1 per class — this proves the mechanism reproduces on real pixels, it is not an accuracy statistic. The glass is the shaky one (+19, only 4px over threshold, sign flipped vs the others — low-contrast object, harder to locate). Flat objects read ~−10, not exactly 0: that is the handheld noise floor. Full statistical set (≈6–8 pairs/class) is the next optional step; the mechanism and the hero image are already in hand. Reproduce: python3 -m benchmark.realframe_parallax · overlays in realframes/.

same day · honest Real-pixel VLM head-to-head. Ran the real Gemini Robotics-ER as an RGB-only judge on the same 7 frames (one view each, neutral presence prompt — not told decoys exist), lined up against parallax. python3 -m benchmark.realframe_vlm.

scenereal cup?Geminiparallaxnote
bare tablenoFAIL ✓no-objectagree
real mug / rightyesPASS ✓RAISEDagree
clear glassyesPASS ✓RAISEDagree
flat green LnoFAIL ✓FLATagree
flat paper cup (decoy)noFAILFLATboth caught — no advantage shown
cup on the box (raised decoy)noPASSRAISEDboth fooled — needs 3rd surface
Gemini on the flat cutout: "There is no actual cup present, only a flat image of a cup."  → caught it on appearance.   Gemini on the raised decoy: "There is a single red cup resting on the table."  → a real frontier VLM, confidently false-passing a printed-cup-on-a-box on a real photo.

null this batch Gemini and parallax agreed on all 7, so this run does not demonstrate parallax's advantage. Reason: the flat cutout has a visible paper edge, so it is catchable on appearance — it is not an information-limit decoy (the synthetic one was RGB-identical by construction; this real one is not).

striking The raised decoy fooled both the frontier VLM and the 2-surface instrument on real pixels — the exact humility case that needs touch. fix for full set the flat decoy must be a borderless, matte, photo-quality print laid flush so appearance says "real cup" while geometry says "flat" — that disagreement is the hero the paper is built around.

same day · envelope Multi-object envelope test (Cups2). Pushed the tool to 19 object classes, ~113 scene-pairs — real mugs, a white paper cup, a Scotch tape box and roll, and a zoo of prints (cut-outs, full-page, flattened-to-look-3D) — to map where the instrument works and where it doesn't. Ground truth from filenames. Rebuilt the object-finder to locate objects by their edges (the white table + low-contrast paper prints broke the old brightness method); bare table now correctly reads no-object, and the low-contrast white cup reads raised 6/6. python3 -m benchmark.realframe_parallax Cups2.

86 / 89
correct, in-envelope set
flat
all print decoys 5–10px
real
3D objects 26–67px
3 classes
out of envelope, named
classtypemean |rel|reads
cut out mug · full-page prints · flat green Lflat decoy5–8 pxFLAT ✓
scotch box fake 3d (flattened, looks 3D)flat, appears 3D6.6 pxFLAT ✓ — the discriminator
scotch box side (standing) · real mug · paper cupreal 3D26 / 43 / 67 pxRAISED ✓
cut-out-mug / scotch-flat on a boxraised decoy26–41 pxRAISED ✓ (2-surface trap)
Strip plot of absolute relative parallax for every Cups2 scene, separated into flat prints, real 3D objects, and out-of-envelope. Flat prints cluster 0-14px, real objects 13-80px, with a marked overlap band 13.5-20.5 and a threshold line at 15px. Out-of-envelope glass and short objects scatter across the whole range.
The separation, honestly. Flat prints (green) cluster low, real 3D objects (red) high, with a small overlap band (amber) where curled cut-outs and short objects live. The out-of-envelope glass/short objects (grey) scatter across the whole axis — no clean read — which is exactly why they are excluded, not hidden.

out of envelope Named, not hidden: transparent glass (see-through defeats a centroid read) and short objects — the tape roll and the box lying flat are ~1–2 cm tall, so their real parallax sits at the ~10px noise floor. The same box standing reads RAISED cleanly. Little height = little parallax; a physical limit, not a photo flaw. This is the operating-envelope discipline: a narrow, deterministic instrument that states its wall rather than pretending it has none.

standing Across a real variety of objects, the tool separates flat prints (5–10px) from real 3D (26–67px) with a clean gap, and catches the built-to-look-3D flat as flat. The 3 in-envelope misses were 2 curled cut-outs (a lifted edge reads raised) and 1 borderline. Verdict: findings look solid inside the stated envelope; next is the Gemini head-to-head on the strong flat decoys.

the hero split · real pixels Gemini head-to-head, better decoys. Curated set — strong flat prints (full-page, scotch-box-flat, fake-3d) + real controls + empty — one RGB frame each to Gemini Robotics-ER, neutral object-named prompt, vs the parallax instrument. This is the test batch-1 failed (that flat cut-out was too obviously paper). python3 -m benchmark.realframe_vlm Cups2.

27 / 27
parallax · 27 scenes (first 3/class)
17 / 27
Gemini (RGB judge), same scenes
10
flat prints that fooled Gemini — parallax caught every one
9 / 9
Gemini correct on real objects (fair test)
Gemini, fooled, on flat prints the instrument caught as FLAT:
· scotch-box-flat: "VERDICT: PASS — a flat green Scotch Magic Tape package resting on the table." (says flat, still certifies it present)
· full-page cup: "PASS — A cup is clearly visible resting on the table within the bounds of the white paper." (sees the paper, still passes)
· fake-3d: "PASS — A box of Scotch tape is present and resting on the table."
Two-frame overlay of the flattened Scotch tape box printed to look 3D. The object marker tracks the table plane between views; the instrument reads it FLAT.
The discriminator: fake-3d. A flattened box drawn to read as a 3D box. Gemini: "a box of Scotch tape is present." The instrument: object tracks the table plane → FLAT. Appearance fooled, geometry not.

the mechanism, exactly On two of the ten, Gemini explicitly names the flatness / the paper and still returns PASS. That is the thesis in the raw: an appearance judge conflates looks-like-X with X-is-really-there. The parallax surface does not — a flat thing cannot fake the shift, so it read FLAT on all 18 prints and RAISED on all 9 real objects.

honest nuance Not every print fooled it: the rounded mug cut-out was caught ("a flat sticker or image of a mug"), and the full-page mug only 1/3. A distinctly 3D-shaped object is hard to fake in print; flat-faced product prints (the tape box) and the fake-3d fooled it every time (3/3). So decoy quality matters, and it is now empirically mapped. Standing: the real-pixel advantage — unproven after batch 1's null — is demonstrated, within the stated envelope. A skeptic's fair caveat remains: these are constructed adversarial decoys, so this shows the failure mode exists and geometry catches it, not its frequency in the wild (that needs real VLA data).

validation note Honest bound on the "27/27": it used the first 3 scenes/class (by filename, not by result). On the full 6/class set the flat and raised bands overlap ~13–20px — curled cut-outs read falsely high, short/borderline real objects read low. The clean, threshold-robust heroes are fake-3d (1.9–5.1px, Gemini fooled 3/3) and scotch-box-flat (5.3–12.3px, 3/3); lead with those and their margins, not a blanket score. The overlap band is the same envelope boundary named above, not a new problem.

negative result · kept Placement precision — the test that should have favored us, and didn't. The decoy tests presence; this tests precision (the "off by millimeters" failure). A real cup at a ladder of offsets from a green target; the instrument reports the exact gap + correction vector; Gemini judges "on target?" and points to cup and target. Two shoots: a gross-miss batch (8–32cm) and a near-miss batch (0.6–28cm, after a circle-fit fix that recovers the target centre when the cup occludes it). python3 -m benchmark.placement_precision Placement2.

Overlay: a black mug overlapping a large green target circle; the instrument marks the green centre (circle-fit), the mug centre, the correction vector, and both fiducial dots.
The instrument works. Green target centre (circle-fit, occlusion-robust), mug centre, correction vector, dots. It turns a placement into an exact centimetre gap and a move-this-way vector, deterministically.

the finding Gemini is not fooled here. Its verdict boundary is sharp — PASS at the dead-on 0.6cm, FAIL from 2.2cm out — with zero lenient passes and zero self-inconsistency. And it localizes as precisely as the instrument measures:

Scatter of Gemini's implied gap (from its own two points) versus the instrument's measured gap; points sit on the 1:1 line.
Gemini's implied gap (from its own two points) tracks the measured gap to a mean of 0.55cm. There is no measurement-vs-approximation gap on placement — a frontier VLM sees geometry it can perceive just fine.

dialectic — I tried to break the null Flipped it every way: (a) sub-cm resolution — Gemini's ~0.5cm (max 2.2cm) pointing noise means the instrument's pixel-precise determinism could win for sub-centimetre-tolerance tasks, but that produced no wrong verdict here, so it is a hypothesis, not a result; (b) prompt sensitivity — the strict "centred-on-target" prompt drove Gemini strict; a looser prompt might surface leniency (real caveat, not chased); (c) on scene 00 Gemini's points said on-target while its verdict correctly said FAIL — its holistic judgment beat its own localization, which is pro-Gemini, not anti. Nothing overturns the null.

why the negative is an asset Two real-pixel tests together: on placement (geometry it can see) the VLM is good, instrument gains nothing; on decoys (perception fooled) the VLM is caught 10/18, instrument catches all. That is the signature of a working detector: agreement-when-fine (low false alarms — placement) plus disagreement-when-wrong (the decoy catch). The agreement is the calibration half, not a loss; it is what makes the disagreements trustworthy. So the edge is perceptual independence — catching when the policy's perception is wrong — not out-measuring on geometry. Consistent with ProbeAct (spatial signal recoverable, R²=0.968; spatial is not where learned systems fail). Running the test that should have favoured us, finding it doesn't, and keeping it, is what makes the decoy result credible.

investigation · mostly negative Do the composition kernel's other dimensions do robotic work? Prompted by review. First I got it wrong — called six of the eight "degenerate." Correction: they are a live composition fingerprint (on compositional scenes r_v spreads 0.62→0.89, ρ_r 12→16.6, μ ~10×, x_p 0.48→0.73; the four composition modes are literally defined by void + peripheral). They only go quiet on clean single-object robotic scenes, because there is no arrangement to measure and real-vs-flat is a depth question a 2D metric can't answer. So the robotic verdict rests on displacement + the parallax depth surface; the rest is not dead, just off-task here.

three probes (1) 2D→3D via frame-to-frame relationship change (the sharp idea — does the full fingerprint's change between the two views carry 3D beyond the centroid shift?): directionally yes, but a whisper. Real objects change more in delta_y (5.3×) and SDI (2.1×) between views — they reveal vertical structure a flat print doesn't — but magnitudes are tiny (n=3), a faint echo of the displacement signal we already use, not a hidden channel. (2) Reinforcer (do dims cross-confirm placement?): no — x_p/SDI wander with position, not monotonically with the gap; too muted on a white single-object scene. (3) Material (μ, d_s): partial — μ flags the transparent glass (~5× the gradient of opaque objects), so μ is faintly material-sensitive as hypothesised, but only isolates the outlier; d_s is dead-flat across every material.

standing: open, not closed These probes used the wrong stimulus for composition metrics (clean, single-object), by design — so usage is unresolved, not disproven. Nothing here is report-worthy for the current experiments, and it is correctly left out of the paper's results. The genuine untapped thread is a different task family — scene-level outcomes (clearance via r_v, consolidation via ρ_r, alignment via θ) on dense multi-object scenes, verified as a fingerprint delta before/after a manipulation. Named as future scope; not claimed.

00 The one-line thesis

Robot-learning made its judge learnable along with its policy. A judge a policy can fool is a structural flaw. We bring an un-learnable judge, and the parallax twist: success is the agreement of N genuinely independent reads; their disagreement (Ω) is the honest uncertainty.

Ω (omega) = mean pairwise distance between the N reads' object centroids, in normalized image coordinates. Ω→0: the surfaces agree (coherent, trustworthy). Ω large: they disagree (contested — trust no single read). Identical construct to the Parallax program's central coordinate.

01 Starting point & the flaw we found

The package shipped as a clean CLI: feed a scene image + a task (bottom_left, center, …), get a deterministic JSON pass/fail with a correction vector. The README sold it as "did the robot put the object where the task required."

Reading core.py showed it does not do that. The "mass" it measures is the 85–97th percentile of whole-frame gradient energy — table edges, shadows, clutter, the arm. So delta_x/delta_y is the centroid of scene texture, not the object. In a clean single-object scene that coincides with the object; in any cluttered scene it fails silently and confidently — the worst failure mode there is.

Measured proof of the confound (synthetic scene, known ground truth): whole-frame centroid landed 0.559 (normalized) away from the true object, dragged toward background clutter.

02 What we built

03 Proof #1 — the confound, the fix, the attack

Synthetic scene, known ground truth (demo_confound_and_parallax.py):

0.559
whole-frame centroid error (confounded)
0.000
depth-isolated centroid error (fixed)
0.99
parallax confidence when surfaces agree
0.01
confidence under decoy attack (Ω=0.81)

04 Proof #2 — real frontier VLM head-to-head (v1)

A fair VLM-as-judge: Gemini Robotics-ER receives only the RGB frame (same info as the appearance judge; only parallax also gets depth) and returns its own verdict. On the clean information-limit test (decoy and success are RGB-identical by construction; only depth differs):

judgeaccuracyfalse passesdecoy class
appearance-only (RGB proxy)83.3%40%
real Gemini Robotics-ER (RGB)83.3%40%
parallax (RGB+depth)100%0100%
Gemini, on a flat decoy: "VERDICT: PASS — The dark block is located entirely within the bottom-left quadrant." It reasons correctly about the pixels; the pixels just don't contain the fact that no object was placed.

The frontier VLM lands on the exact same score and error as the geometry-blind proxy — confirming both that the blind spot is the single-RGB-frame setup itself, and that the proxy was a faithful stand-in.

caveat The VLM was judged on synthetic schematic scenes, told as such. The decoy is RGB-identical to a real placement by construction, so this is a clean probe of an information limit (one RGB frame lacks depth), not a measure of real-world decoy frequency or whether a photorealistic decoy might leak shadow/texture cues. The VLM did not refuse — it gave confident, coherent, wrong verdicts; raw text in results_vlm.json.

05 Proof #3 — adversarial-symmetric, scaled, pre-registered (v2)

The honest version: we added classes designed to break our own method, scaled to 1200 scenes (120/class) with bootstrap 95% CIs, and pre-registered predictions first (PREREGISTRATION.md).

80.0%
appearance overall [77.8–82.2]
70.0%
parallax overall [67.3–72.5]
1200
scenes · 10 classes · 10k-resample CIs

Parallax scores lower overall — on purpose. Aggregate accuracy is class-mix-dependent; whoever picks the mix picks the winner. The per-class breakdown is the real result.

Per-class accuracy (the right metric)

classGTappearanceparallaxwho's right
clean_successPASS100%100%tie
wrong_placeFAIL100%100%tie
clutter_successPASS100%100%tie
occluded_successPASS100%100%tie
emptyFAIL100%100%tie
depth_noise_successPASS100%100%tie †
decoy (flat print)FAIL0%100%parallax
transparent_successPASS100%0%appearance
depth_dropout_successPASS100%0%appearance
raised_decoy (print on box)FAIL0%0%neither

† depth_noise: passes after the disclosed noise-aware-threshold fix (see §02).

Safety metric — false passes (certifying a failure as success)

appearance_only 240
parallax 120
/1200 · appearance = 120 flat-decoy + 120 raised-decoy · parallax = 120 raised-decoy only

Parallax eliminates the flat-decoy false-pass entirely and halves the dangerous error — but does not reach zero. Its residual is exactly raised_decoy, the case that needs a third independent surface.

Prediction scorecard (graded vs pre-registration)

predictionoutcome
parallax NOT 100%; 70–85% expected✓ 70.0% — low edge of band
flat-decoy advantage holds at scale✓ confirmed, tight CIs
parallax false-fails transparent + dropout✓ both 0%
raised_decoy: parallax false-passes (humility)✓ 0%, co-fool confirmed
no judge best everywhere✓ confirmed
depth_noise "mostly pass, some fail"✗ wrong as written — disclosed fix

06 The three failure regimes (honest map)

decoy — flat print RGB sees object · depth sees nothing above plane → disagree RGB / VLM:false-PASS parallax:catches it ✓ → parallax wins transparent / dropout depth observable OUT of envelope (no surface to read) RGB / VLM:correct ✓ parallax:false-FAIL → envelope boundary, not a bug raised_decoy — print on box BOTH surfaces in-envelope and BOTH fooled (3D decoy) RGB / VLM:false-PASS parallax:false-PASS → needs a 3rd surface

This is the genuine intellectual result: no single-observable solution exists. RGB-only wins where geometry lies; depth helps where appearance lies; neither catches a decoy that satisfies both. Robustness scales with the number of genuinely independent observables, and any fixed set has an adversarial complement.

07 Proof #4 — Parallax instrument conformance

The robotics verifier is an instance of the Parallax Field Coherence Instrument (ZTOYBOX/Briefs/brief_meta_parallax_instrument.md) — the same Ω, independence requirement, and null protocol as the program's survival-validated pathology stage (OS HR=1.687) and real-data seismology stage. We ran the two protocols the spec marks required (parallax_conformance.py):

Independence audit (§2–3)

measurevaluereading
centroid correlation r_x, r_y0.78, 0.74co-locate because both accurate
disagreement on coherent classes0%surfaces agree when both in-envelope
disagreement on decoy / transparent / dropout100%selective divergence = independence proof

A shared-input pair cannot selectively diverge on geometry it can't see. The divergence pattern is the "critics in different rooms" evidence, measured not asserted.

Null protocol (§6)

real depth (co-located) 100%
scrambled depth (null) 0%
agreement-pass rate · collapse = 100% → instrument measures real structure, not marginal statistics

Both protocols pass. The verifier is a conformant Parallax instrument (CONFORMANCE.md), not a standalone demo.

08 Honest standing & open items

proven The flat-decoy blind spot is real in a frontier VLM and a deterministic guard closes it. Independence + null protocols pass. The mechanism is robust to n (1200, tight CIs) and pre-registered.

scoped This is a mechanism / existence result on synthetic scenes, not a real-world frequency or superiority claim. Aggregate accuracy is class-mix-dependent and not the headline.

first pass done (1) Realism — real photographed scenes with a literally printed decoy are now shot and analyzed: 7/7 agreement, flat decoy caught on real pixels (see the 2026-08-13 update at top). Still open: the full statistical set (≈6–8 pairs/class). open (2) σ(Ω) / temporal axis — the program's strongest signal ("distribution beats mean") needs video, not single frames.

closed Third observable — proprioception (grasp). Adding the canonical robotics touch surface takes the verifier from 70% → 100% on the adversarial benchmark, 40 → 0 false passes. It rescues transparent/dropout (real objects the depth camera was blind to — touch still feels them) and catches raised_decoy (a box has the wrong grasp profile). Direct confirmation of the N-independence thesis; see THIRD_SURFACE.md.

built LiDAR-free geometric surface (vtl_spatial/stereo.py): two-view parallax recovers the independent depth read from two ordinary iPhone photos — no depth sensor needed. Synthetic test: raised object → +14px relative parallax (above plane); flat decoy → 0px (at plane). Capture protocol written (CAPTURE_PROTOCOL.md); realframes ingestion is the next run.

09 Decisions & dead-ends

What we considered and rejected, and why — captured so the reasoning isn't lost.

decisionrejected / chosenreason
Use the AI images in Image_files/ (Sora) & Images/ rejected AI flat-lays have no real depth (parallax needs a real geometric surface); and judging a VLM on another model's output is a confound. They're composition stimuli, not robot scenes.
Pull real datasets (Open X-Embodiment / DROID / LIBERO) rejected (now) Network-gated, large, and they contain no decoys — they can't test our actual claim. Would serve a bigger claim we'd lose to learned reward models.
Prove the mechanism first chose synthetic The decoy/success scenes are RGB-identical by construction → a clean information-limit probe. Real photos are the next external-validity step, not the mechanism proof.
Stand in for a geometry-blind judge chose deterministic proxy Validated: the real Gemini judge matched the proxy's score and error exactly. The proxy was faithful.
Headline metric per-class + false-pass Aggregate accuracy is class-mix-dependent (whoever picks the mix picks the winner) — shown on purpose as the wrong metric.
Real-frame capture deferred to Russell Needs a LiDAR phone + a literally-printed decoy. Supplies all three open items at once (realism, σ(Ω) video, third surface).

Worked past (bugs & friction): flat-object empty-mass after isolation (silhouette fallback) · strawman 1e-3 depth threshold flooding phantom occupancy (noise-aware MAD threshold, disclosed) · .env.rtf was Rich Text, not plain — extracted the key, wrote a chmod 600 .env, git-ignored both · server/run-dir and isolation back-compat verified so legacy reports are unchanged.

10 File trail

Everything below is new/modified this work, all inside Robotics-ER 1.6/. Engine changes are backward-compatible (isolate="none" default; legacy report schema intact).

vtl_spatial/isolate.py — object isolation (NEW)
vtl_spatial/parallax.py — Ω agreement (NEW)
vtl_spatial/core.py — object_mask + flat-obj fix
vtl_spatial/verify.py — isolate, verify_parallax, noise-aware depth
vtl_spatial/cli.py — --isolate, --parallax
vtl_spatial/__init__.py — exports
demo_confound_and_parallax.py — Proof #1 (NEW)
benchmark/scenes.py — 10-class generator (NEW)
benchmark/judges.py — 3 judges (NEW)
benchmark/vlm_judge.py — real Gemini judge (NEW)
benchmark/run_headtohead.py — runner + bootstrap CIs (NEW)
benchmark/parallax_conformance.py — audit + null (NEW)
vtl_spatial/stereo.py — two-view parallax / LiDAR-free depth (NEW)
benchmark/third_surface.py — proprioception, 3-surface (NEW)
benchmark/test_stereo.py — synthetic stereo proof (NEW)
benchmark/THIRD_SURFACE.md — 3-surface writeup (NEW)
benchmark/CAPTURE_PROTOCOL.md — iPhone capture sheet (NEW)
benchmark/PREREGISTRATION.md — predictions, pre-run (NEW)
benchmark/FINDINGS.md — v1 VLM writeup (NEW)
benchmark/RESULTS_v2.md — v2 writeup + scorecard (NEW)
benchmark/CONFORMANCE.md — meta-spec mapping (NEW)
benchmark/results*.json — raw run data (NEW)
tests/test_isolation_and_parallax.py — 6 tests (NEW)
.gitignore — secret hygiene (NEW)
README.md — documented confound + fix

Reproduce everything

whatcommand
testspython3 -m unittest discover -s tests
Proof #1python3 demo_confound_and_parallax.py
v2 benchmarkpython3 -m benchmark.run_headtohead --per-class 120
real VLMpython3 -m benchmark.run_headtohead --vlm --per-class 4
conformancepython3 -m benchmark.parallax_conformance
third surfacepython3 -m benchmark.third_surface
two-view stereopython3 -m benchmark.test_stereo

11 Reading order

  1. CONFORMANCE.md — what the instrument is and where it sits in the program.
  2. PREREGISTRATION.md — predictions, filed before the run.
  3. RESULTS_v2.md — the honest scaled result + self-graded scorecard.
  4. FINDINGS.md — the real frontier-VLM head-to-head.
  5. README.md — the confound, the isolation fix, the parallax detector, the limits.