Parallax Metrology
VTL Spatial · embodied verification
Technical white paper · Spatial AI verification

Does it look right, or is it right?

On real photographs, a deterministic two-view geometric check separates real 3D objects from flat printed decoys within a stated envelope — and catches ten flat decoys a frontier VLM certified as real objects present. It is a spatial GPS for embodied AI: it doesn't replace the VLM judge or the policy, it runs beside them, sub-second and un-gameable, and flags where appearance and geometry disagree — the candidate poisoned label. What follows is that result, and an honest map of what it catches, what it misses, and where a partner takes it next.

Fig. 0 — the instrument on a real scene · task: place mass bottom-left
A cluttered desk photographed top-down. A red marker labeled 'measured' sits near the coffee cup at frame center; a green marker labeled 'target' sits in the bottom-left corner; a blue correction vector connects them. A banner reads FAIL score=0.425.
The whole thesis in one frame. A person asked for the objects bottom-left; perception agrees they read as lower-left. The kernel measures the actual centroid of visual mass — it lands up by the coffee cup, well short of the bottom-left target (measured vs target). The blue arrow is the correction vector — direction and magnitude, not just a verdict: move Δx −0.22, Δy +0.23. The weighted compliance gap, 0.425, far exceeds the 0.18 tolerance → FAIL.
Specimen readout · Fig. 0actual engine output · schema vtl-spatial-compliance/v1
Δx measured
−0.181 / tgt −0.40
Δy measured
+0.118 / tgt +0.35
void ratio rv
0.881
target mass
0.589 < 0.62
outside leakage
0.411 > 0.38
verdict
FAIL · 0.425
Named failure modes: centroid_miss · target_under_occupied · outside_leakage · under_packed. Every field is a reproducible measurement, not a model opinion.
Context · why a non-learnable cross-check, and why now

The field is wiring a foolable judge into the data engine

To scale training data for robot policies (VLAs), teams increasingly use a large VLM as an automated judge that scores whether an episode succeeded — shipped, e.g., in one production toolchain (VLM-as-judge episode scoring whose reward labels feed VLA fine-tuning). But the field documents how unreliable that judge is: on the RoboTwin 2.0 manipulation benchmark, a VLM success/failure observer scored precision 0.208 — of every five episodes it flags as successes, four are not — badly calibrated in both directions (RoboTwin 2.0). And robotics groups name the exact failure this work is about — VLA policies "fail silently: a collapsing rollout looks much like one making clean progress" (ValueFormer) — a live enough bottleneck that a 2026 cluster races to detect it (VLA-FAIL, SAFECAST). One study frames the mode in this paper's own terms — a Berkeley observability analysis of induced (not constructed) false successes finds "a peg that looks seated but is not" leaves almost no trace, corroborating that the silent-failure class occurs beyond printed decoys (observability study); notably, it also finds much of that signal is recoverable from proprioception — a point returned to in the ask.

Every one of those fixes is still learned — a better-trained judge is still a judge a policy can learn to fool. This work moves on a different axis: a deterministic, non-learnable geometric cross-check with no learned surface to attack. It doesn't try to win the judging contest; it grounds it. A spatial-GPS anchor that partners with the learned stack, where the two disagreeing is itself the signal. Independently, from image composition rather than robotics, it converges on the field's own stated bottleneck: silent failures poison the training labels.

I

The gap, and where the instrument came from

Lineage · the composition origin — context, not the proof (the proof is Movement III)

Robots can see and robots can plan. What no standard pipeline instruments is whether the spatial intent of a natural-language instruction became a geometric fact. Semantic classifiers answer "does this look correct?" Point clouds handle collision geometry. Neither produces a scalar compliance gap when an instruction is satisfied perceptually but not geometrically.

The instrument: displacement, plus an independent depth cross-check

The origin is an 8-dimensional composition kernel — a deterministic map from a frame to geometric field coordinates. As an aggregate fingerprint it carries real signal on the scenes it was built for: across compositional images, void ratio, packing, mass intensity and peripheral pull vary substantially and separate one arrangement from another (no single dimension is the story — the pattern across the vector, and the comparison between scenes, is). But it is worth being exact about which of it does the robotic work, because the robot task asks a narrower question. Real-vs-flat is a depth question — and a 2D composition metric cannot see depth at all; moreover, an isolated object on an empty table has no compositional arrangement for those dimensions to register, so on these tabletop scenes they sit near-constant across real and flat alike. So the robotic verdict rests on displacement (Δx, Δy) — did the mass land where instructed — plus a surface the kernel does not compute at all: an independent two-view parallax depth read. The honest description of the robotic instrument is displacement plus a depth cross-check; the kernel's other dimensions are live on the composition problem they came from, not dead weight — they simply do not answer the depth question the robot task poses.

K(I) = [ Δx, Δy, rv, ρr, μ, xp, θ, ds ]  — a composition fingerprint

Δx · Δy
Placement offset
Normalized centroid displacement of visual mass from a task-defined origin.
rv
Void ratio
Fraction of the frame below the gradient-activity threshold — open vs. congested.
ρr
Packing density
Clustering of occupied regions relative to their convex hull.
xp
Peripheral pull
Proportion of mass within the outer 15% frame band — edge-drawn vs. central.
μ
Mass intensity
Mean gradient magnitude across the active field — how strongly the occupied region registers.
ds
Structural thickness
Skeleton-normalized average material thickness — thin vs. volumetrically heavy.
θ
Orientation stability
Entropy-based directional coherence — aligned vs. chaotic structure.

The eight are hand-designed features, not learned; fidelity as a proxy for "the object went where instructed" is measured here on constructed inputs, with validation on real rollouts still open. Stated positively: this is a spatial field, not a displacement gauge — its full multi-dimensional range is demonstrated on compositional scenes (Fig. 2 — void, peripheral pull, and packing all carry signal). The single-object robot demonstration simply does not exercise it, and that range applies naturally to scene-level robot verification — did the workspace get cleared, consolidated, aligned — the instrument's growth path, demonstrated on compositions, not yet on robots.

Where it came from: instruction intensity barely moves geometry

The kernel was developed on a different problem — spatial behavior in text-to-image generation. Across 6,000+ generated images on 10 platforms, spatial prompt intensity explains 0.0–0.1% of compositional-displacement variance; if instruction perfectly drove placement, the slope would be +1.000, and it isn't close. It is a real, independent finding about generative-model controllability, and the failure it names — a model satisfying a spatial instruction perceptually without executing it geometrically — is what pointed toward the robotics work.

Stated plainly, so it is not load-bearing A generated image has no physical object that was or wasn't placed — only pixels a diffusion model produced. So the compliance gap measured there is not the same object as the gap between an instruction and a gripper's outcome, and this paper does not claim the transposition proves anything. The composition study is the origin; the evidence that the instrument works on physical outcomes is the real-pixel robotics result in Movement III, and that stands on its own.
6,000+
generated images · 10 platforms
0.0–0.1%
displacement variance explained by prompt intensity
±0.02
measurement SD · 55+ regenerations
Slope of geometric response to instruction intensityperfect compliance = +1.000
perfect
+1.000
Sora
+0.019
MidJourney
+0.024
OpenArt
−0.050
Engines overshoot their neutral boundary at mild instruction, then retract under escalation. OpenArt's slope is negative — harder instructions produce marginally less displacement. The controllable range between the envelope boundary and perfect compliance is never occupied.

The mechanism is a constraint substitution: models satisfy a spatial instruction by swapping in a perceptual strategy — clustering, dispersion, symmetry, partial placement — none of which guarantees the mass actually moved. In four hand-audited real scenes, humans said success, Gemini said success, and the kernel said failure. The kernel makes that gap a number.

Fig. 1 — what perception reads vs. what the field measures
Four panels of a woman reading by a window: the original photo; a gradient mass-and-skeleton map; the same image marked with a measured centroid at Δx = −0.143; and a radial safety overlay.
A "left-weighted" photo whose mass is centered. The figure and window read left, but the book (sharp edges, center), hair volume (right), and blinds counterbalance into an L-shaped mass around the frame midpoint — Δx = −0.143, not the strong leftward value perception predicts. The kernel does a global accounting of all visual mass; it does not privilege the dominant cluster.
Fig. 2 — four modes of spatial non-compliance
Infographic titled Four Modes of Spatial Non-Compliance with four labeled panels: Cluster Bias, Peripheral Dispersion, Symmetry Without Occupancy, and Partial Placement Override. Each shows object mass, a measured centroid, a target location, and a correction vector; each is labeled Perception PASS, Geometry FAIL.
Perception PASS, geometry FAIL — four distinct failure fingerprints. Cluster bias, peripheral dispersion, hollow symmetry, and partial placement each satisfy a human read while leaving the centroid off-target. The mode determines the correction strategy — the same kernel, unmodified, produces a different correction vector for each. (This legend — object mass, measured centroid, target, correction vector — is reused in every chart below.)
Fig. 3 — the same failure, in a robotics frame
A top-down bin-sorting scene captioned 'Place all tools in the bottom-left bin.' A vision system check marks the tools as being in the lower-left area, while the kernel marks the measured centroid still near frame center and reports a compliance gap and correction vector.
"Place all tools in the bottom-left bin." The vision system checks and confirms: objects are in the lower-left area. It misses what the kernel measures — the centroid is still near center, so the task is not complete. The kernel returns a quantitative verdict: gap 0.139, correction Δx −0.04, Δy +0.12. This is the composition finding above, transposed onto exactly the verification step a manipulation pipeline has to get right.
II

From a confounded demo to a conformant instrument

Additional work · what was wrong, what we built, what we proved

The package shipped as a clean CLI that claimed to answer "did the robot put the object where the task required." Reading the source showed it did not. Fixing that — honestly, adversarially, and against a real frontier VLM — is the body of this work.

The confound: it was measuring the scene, not the object

The kernel's "mass" was the 85–97th percentile of whole-frame gradient energy: table edges, shadows, clutter, the robot arm. So Δx/Δy was the centroid of scene texture, not the manipuland. In a clean single-object scene the two coincide; in any cluttered real scene the verifier fails silently and confidently — the worst failure mode there is. On a synthetic scene with known ground truth, the whole-frame centroid landed 0.559 (normalized) away from the true object, dragged toward background clutter.

The fix: isolate the object before the kernel runs — without a learned judge

Object isolation runs ahead of the kernel, in order of trust: an explicit segmenter mask › depth-plane removal › largest connected component. Perception may be learned; the verdict stays deterministic. This is the architecture the whole program implies — learned eyes, un-learned judge. Depth-plane isolation recovered the true centroid from 0.559 error to 0.000.

Object-centroid error — normalized image unitssynthetic scene, known ground truth
whole-frame
0.559
depth-isolated
0.000
The confound (red) is not a small bias — it is more than half the frame. Isolation (green) removes it entirely.

The parallax turn: success as the agreement of independent reads

A single deterministic check is brittle; a single learned judge can be fooled by the very policy it grades. The fix reframes success itself: it is the agreement of N genuinely independent reads of one outcome, and their disagreement — Ω, the mean pairwise distance between the reads' object centroids — is the honest confidence.

Ω → 0: the surfaces agree — coherent, trustworthy. Ω large: the observables disagree — the outcome is contested, not passed. Trust no single read.

— the same coordinate the Parallax program uses in seismology and pathology, transposed to embodied verification

Independence is load-bearing: N reads that are all transforms of the same RGB surface measure transform diversity, not physical agreement. Each read declares its observable surfacergb, depth, proprio, time — the audit counts distinct surfaces, and a single-surface verdict is capped at 0.5 confidence. Under a decoy attack (a flat printed object that satisfies appearance but has no geometry), Ω spikes to 0.81 and confidence collapses to 0.01: correctly contested, where a lone VLM judge is fooled.

Proof: a frontier embodied VLM scores exactly like a geometry-blind proxy

We ran the parallax detector against two comparators on a labeled suite: an appearance-only proxy (RGB, geometry-blind) and the real Gemini Robotics-ER 1.6, given the same single RGB frame and a neutral judge prompt. The result is the headline.

Head-to-head · labeled suite · overall accuracyn = 24 · 6 classes · false passes at right
appearance
83.3%
Gemini ER 1.6
83.3%
parallax
100%
RGB-only judges  ·  parallax (RGB + depth)  ·  false passes: appearance 4 · Gemini 4 · parallax 0

The frontier embodied-reasoning VLM lands on the exact same score and the exact same error as the geometry-blind proxy: both false-pass every decoy. That dual result is the point — it confirms both that the proxy was a faithful stand-in, and that the blind spot is not a weakness of one model but of the single-RGB-frame setup itself.

"VERDICT: PASS — The dark block is located entirely within the bottom-left quadrant."

— Gemini Robotics-ER 1.6, on a flat printed decoy. It reasons correctly about the pixels; the pixels do not contain the fact that no object was placed.

The honest version: adversarial-symmetric, pre-registered, scaled to 1,200

The first suite only contained classes that break RGB judges — a cherry-pick. We corrected it: added classes designed to break our own depth surface, scaled to 1,200 scenes (120/class) with bootstrap 95% CIs, and pre-registered every prediction — including predicted failures of our method — before running.

Per-failure-mode accuracy · 1,200 scenes · 10 classesthe right metric
ClassGround truthAppearance (RGB)Parallax (RGB+depth)Who is right
clean_successPASS100%100%tie
wrong_placeFAIL100%100%tie
clutter_successPASS100%100%tie
occluded_successPASS100%100%tie
emptyFAIL100%100%tie
depth_noise_successPASS100%100%tie †
decoy (flat print)FAIL0%100%parallax
transparent_success (glass)PASS100%0%appearance
depth_dropout_successPASS100%0%appearance
raised_decoy (print on box)FAIL0%0%neither
Overall accuracy: appearance 80.0% [77.8–82.2] · parallax 70.0% [67.3–72.5]. Parallax scores lower overall — on purpose. Aggregate accuracy is class-mix-dependent; whoever picks the mix picks the winner.  † after a disclosed noise-aware-threshold fix.

No judge dominates, and that is the honest conclusion. RGB wins where depth has no signal (transparent, dropout). Parallax wins on the flat decoy. And a printed photo mounted on a low box — raised_decoy — fools both surfaces at once. The safety metric that actually matters is the false pass: certifying a failed outcome as success poisons a training set and green-lights a bad action.

Read these synthetic per-class numbers as an upper bound. They come from a noiseless generator, where the flat-vs-raised separation is cleaner than any real sensor delivers — the same caveat the observability study makes about its own simulated separability. The load-bearing evidence is not this table; it is the real-pixel margin in Movement III (flat prints ~5–10 px vs real objects ~26–67 px, 86/89 in-envelope), measured on photographs with real noise.

Appearance-only false passes
240
/ 1,200 · flat decoy + raised decoy
Parallax false passes
120
/ 1,200 · raised decoy only

Parallax eliminates the flat-decoy false pass entirely and halves the dangerous error — but does not reach zero. Read the two numbers precisely: it removes the camera-closable half (120 flat-decoy false passes → 0), and the entire residual (120) is raised_decoy — the one failure mode cameras cannot close, because both camera surfaces are fooled at once. So the residual is not a loose end; it is the exact size of what a third, non-camera surface must catch. Two independent surfaces are a floor, not a ceiling — and the 120 you can't close with cameras is precisely why the robotics partner is needed.

Conformance — not an ad-hoc tool The verifier is run as an instance of the Parallax Field-Coherence Instrument, and inherits its required protocols. Independence audit: the surfaces agree on coherent scenes but selectively diverge on exactly the decoy / transparent / dropout classes — a shared-input pair cannot selectively diverge on geometry it cannot see, so the divergence pattern is the independence proof. Null protocol: scramble the depth structure and the agreement collapses from 100% to 0% — the instrument measures real co-location, not marginal statistics. Both run and pass.

Two independent surfaces are a floor, not a ceiling — and the raised_decoy residual points at exactly what a third one has to be: a channel the cameras don't share. That is where a solo synthetic study becomes a collaboration, and it is taken up as the next step, not claimed as a result (below).

This verifier is the embodied-verification stage of a broader deterministic-measurement program — parallaxmetrology.com — but here it stands only on its own robotics evidence, above.

III

Real pixels: the mechanism reproduces, and a frontier VLM is caught

The synthetic proof, earned on real photographs — with the boundary stated

The capture happened. Real objects, photographed on an ordinary phone from two ~10 cm-offset positions — no depth sensor; the second view is the geometry. Run across 19 object types (real mugs, a paper cup, a Scotch-tape box and roll, and a range of printed decoys), the synthetic mechanism holds on real pixels.

Real 3D vs. flat print — a clean split, and an honest boundary

Within a stated operating envelope, the instrument separates real 3D objects from flat prints with a wide margin: flat prints register ~5–10 px of relative parallax, real objects ~26–67 px — 86/89 in-envelope. The envelope is named, not hidden: the instrument measures height above the table plane, so transparent objects (a camera sees through them) and objects too short to parallax (a tape roll, a box lying flat) are out of it. A narrow instrument with a stated boundary, rather than a vague one. On transparency specifically, one coherent reading resolves what look like three: the instrument's intended behavior is to exclude glass as out-of-envelope; the varied empirical reads across tests (an anomalous single-frame raise, a synthetic false-fail, real detect-failures) are simply the noisy edge of that boundary, not conflicting results — a transparent object gives the depth surface nothing stable to measure, by design.

The head-to-head: 10 flat decoys a frontier VLM certified as real

On a curated set of the strong flat decoys plus real controls, one RGB frame each went to Gemini Robotics-ER (neutral, object-named prompt) alongside the instrument. Gemini was fooled by 10 of 18 flat prints — and the instrument read every one as flat. The cleanest, threshold-robust cases are the flattened-box fake-3d (parallax 1.9–5.1 px, Gemini fooled 3/3) and the flat tape-box print (5.3–12.3 px, 3/3): comfortable geometric margins, on prints a frontier VLM certified as real objects present.

"PASS — a flat green Scotch Magic Tape package resting on the table."  ·  "PASS — a cup is clearly visible resting on the table within the bounds of the white paper."

— Gemini, on two flat prints the instrument caught. It names the flatness / the paper, and certifies the object present anyway.

Two distinct failures live here, and they should not be conflated. When the decoy is RGB-identical to a real placement (the synthetic case, by construction), no RGB-only judge can tell — that is an information limit, and the second observable is genuinely necessary. But on the real prints, Gemini named the flatness — "a flat green package," "within the bounds of the white paper" — and passed anyway. There the information was in the frame and the model extracted it; its success-verdict simply ignored its own perception. That is not an information limit, it is a judgment failure decoupled from perception — arguably the stronger finding, because you cannot fix it by adding information. In both, a flat thing cannot fake the parallax shift, so the geometric check does not make the mistake; but the reason it wins differs — necessity of a second observable in one case, deterministic independence from a fooled judgment in the other.

Fig. 4 — the discriminator, on real pixels · "fake-3d"
Two-frame overlay of a flattened Scotch tape box printed to read as three-dimensional; the object marker tracks the table plane between the two views, and the instrument reads it FLAT.
A flattened box, drawn to look 3D. Gemini: "a box of Scotch tape is present and resting on the table." The instrument: the object tracks the table plane between the two views → FLAT. Appearance fooled; geometry not.
Honest bounds These are constructed adversarial decoys — this shows the failure mode exists and geometry catches it, not its frequency in the wild. The head-to-head used the first three scenes per class; on the full set the flat and raised bands overlap ~13–20 px (curled cut-outs, short objects) — which is why the claim rests on the two robust heroes' margins above, not a blanket per-class score. And not every print fools the VLM — a rounded mug cut-out was caught ("a flat sticker or image of a mug"); flat-faced product prints and the built-to-look-3D fooled it every time. The clean, robust demonstrations are the fake-3d and the flat tape-box print.

Couldn't a better-trained judge just catch these? — two reasons it can't

The obvious objection: train a learned reward model on enough real data and it catches the decoys too. Two distinct answers, kept apart:

Placement precision — where the two agree, and why that is the other half

The decoy tests presence. We also ran the fake-free test of precision — a real cup at a ladder of offsets (0.6 to 28 cm) from a target — precisely because it should favor a measuring instrument over a VLM. It did not, and that is reported plainly. Gemini localizes as precisely as the instrument measures (its gap, read from its own two points, tracks the measured gap to a mean of 0.55 cm), and its verdict is sharp — it passes the dead-on placement and correctly fails everything from 2.2 cm out, with no lenient passes. On geometry it can perceive, a frontier VLM is a competent spatial judge, and the instrument holds no advantage.

That agreement is not a surprise — spatial localization is something embodied systems already do well — and it is not a consolation prize either. It is the calibration half of a working detector. The product is "run this beside the learned judge and flag the disagreements as candidate poisoned labels," and that requires two things: agreement when the outcome is fine (a low false-alarm rate) and disagreement when it is wrong (a true catch). Placement supplies the first; the decoys supply the second.

Agree where the judge is right, disagree only where it is wrong Put the two real-pixel tests together and the behavior is one coherent thing, not two results. On placement (geometry the VLM can see) the two agree — validating the instrument and showing it will not cry wolf on ordinary, correct placements. On decoys (perception fooled) they disagree 10/18 — the instrument catching what the judge misses. A cross-check that disagreed everywhere would be noise; one that agreed everywhere would be redundant. This one disagrees selectively, exactly when perception is wrong — which is the whole value, and the reason its disagreements can be trusted as signal. The edge, stated precisely, is perceptual independence, not out-measuring geometry; consistent with ProbeAct's finding that spatial reasoning is not where learned systems fail. (Honest bounds: the placement verdict was elicited with a strict, centered-on-target prompt; and both sides here read geometry, so the agreement validates accuracy, not the RGB-vs-depth independence, which is a separate pairing.)

The third surface — the natural place a partner comes in

The two-surface residual (raised_decoy) is not a dead end; it is a specification. Robustness scales with the number of genuinely independent observables, and the missing one is the channel a camera cannot spoof: proprioception — did the gripper actually close on real mass of the expected profile? Touch, not vision. In a controlled study it closes every remaining hole: a blind depth surface (glass, occlusion) no longer vetoes an object the gripper can feel, and a raised decoy fails because a box has the wrong grasp profile.

Grounded by the win, honest about the next step The win this paper stands on is the one above: on a pre-registered 1,200-scene benchmark, the deterministic cross-check eliminates the flat-decoy false pass and halves the dangerous error, using cameras alone. The third surface is the next step, not one of those numbers — and the controlled study that closes the residual models the grasp signal from each scene's true height, so it demonstrates the architecture rather than a live gripper. The real signal comes from real hardware. That is precisely the seam where this meets an embodied-AI team: the cross-check is built and proven on cameras; the third surface is a channel a robotics partner already has on the arm.

Where this sits, and what is actually new

Adjacent work — conceding the commodity floor, and naming the unoccupied combination

Geometry as a second signal for embodied AI is not new. The parts of this that are common should be conceded fast, because the novelty is a specific combination, not "add geometry."

The commodity floor. Using 3D structure to help a VLA is occupied. VeriSpace is a 3D-aware verifier that re-ranks candidate actions before execution; multi-view geometric consistency — the parallax mechanism used here — is classical structure-from-motion, table stakes rather than a contribution. Neither is claimed. This work does a different job (audit a completed outcome for a poisoned label, not select an action) with a different kind of check (deterministic, not learned).

What is unoccupied is the combination: a check that is deterministic (no learned surface for a policy to exploit), independent of the policy (so it can catch the policy's own perceptual errors), applied after the fact as a label-quality filter, where the disagreement between it and the learned judge is the signal.

The sharpest neighbor — ProbeAct, and a clean division of labor ProbeAct reads the correct 3D object position straight out of an OpenVLA policy's own hidden states (R²=0.968 at layer 8; probe error 6.9 cm while the drifting action endpoint is 23.6 cm) and concludes "external sensors are redundant." Read carefully, it strengthens this work rather than subsuming it. ProbeAct's probe lives inside the policy being verified, so it catches the "Phantom Grasp" — perception right, motor wrong — and structurally cannot catch perception itself being wrong: a policy fooled by a flat print encodes "object present" in the very hidden states the probe reads. That is the flat-decoy case here, exactly. The two cover disjoint failure modes — ProbeAct, motor drift when perception is right; this, perceptual lies, because it is independent of the policy. And ProbeAct's own failure detector — a deterministic kinematic check of whether the object rose synchronously with the gripperis the proprioceptive third surface this paper names as the next step. (ProbeAct is simulation-only, sim-to-real flagged as future work; the result here runs on real photographed pixels.) The refinement it forces is worth stating precisely: the requirement is not an external observable, it is an observable independent of the policy.
IV

The larger idea: instructions as measurement spaces

The aspiration · from a verifier to a spatial-intent compiler

The current kernel measures displacement from the frame center. But the center is the wrong origin for almost every spatial instruction. "Place it in the corner" means the corner is the zero point — compliance should be distance from the intent, not from the frame. The delta has to move with the task.

That points past a verifier toward a different architecture. A natural-language instruction does not just define a target state — it initializes the coordinate system. Before the kernels run, the instruction is parsed to establish three things:

The kernels then read the field within that instruction-defined space. The same physical scene produces different vectors under different instructions — which is correct: compliance is always relative to intent. Add the Z axis — depth as a first-class dimension, not an edge case — and you have genuine 3D coordinates tied to intent.

Spatial instructions compile into measurement spaces, and execution is verified within them. The kernel stops being an approximation engine and becomes a coordinate engine — one component of a spatial-intent compiler.

— the design note this work opens onto

That is a larger and more consequential idea than a success detector, and the verifier is its first working component: the piece that measures, deterministically, whether intent became geometry.

What would break this

Where the claim is a theorem, where it is falsifiable, and the one result that would kill it

The central claim — robustness scales with the number of genuinely independent observables — has to be pinned down, because part of it is true by construction and part is empirical, and conflating them is what makes it feel slippery. Kept apart:

The one result that would kill the practical claim If a learned RGB-only judge, given enough real rollout data, matched the parallax instrument on the decoys that actually occur in the wild, the practical independence advantage would be dead — it would mean real silent failures are mostly RGB-detectable, and the second observable buys little. That is the honest falsification condition, and it is exactly what real VLA data settles. The information theorem would survive; the value proposition would not.
The ask

Run it beside your judge. Where they disagree is your candidate poisoned label.

The concrete niche is a label-quality filter in a VLA data engine: run this deterministic geometric cross-check alongside the VLM judge already scoring your episodes, and flag disagreement for review. It is cheap (sub-second, no GPU), reproducible, and has no learned surface for a policy to game — so it catches the confident false pass a better-trained judge would still wave through. Not a replacement for the learned stack: the ground-truth anchor beside it.

Two honesties that belong in the ask, not a footnote. First, the disagreement signal is trustworthy only inside the height-above-plane envelope. On transparent or camera-occluded objects the instrument disagrees when it is wrong — there it should defer, not fire. A deployable filter runs only where its observable is in-envelope. Second, mind the evidence tiers. The camera-only cross-check that halves the dangerous error is built and proven on real pixels; the residual it does not close (a 3D decoy that parallaxes like a real object) is closed only by a third surface — proprioception — which is modeled here, not built. Independent work (ProbeAct; the Berkeley observability study) suggests the touch/proprioception channel does much of the real recovery — corroboration for the multi-surface architecture, and the reason the third surface is the collaboration, not a shipped feature. It also invites the sharp question: if proprioception recovers most of it, is the built camera surface redundant? The honest parry: parallax is not the highest-recall channel, it is the independent one, and it is distinct in two ways proprioception is not. It is non-contact and works from a recorded video alone — so it audits logged episodes for label quality where no proprioception was captured (the data-engine case); and it is an independent geometric read inside the vision modality, so it catches a perceptual lie the learned vision judge shares and touch would only reach after the robot has already acted on the bad label. Where proprioception is logged, the two overlap on contact failures — conceded; where it isn't, or where the goal is offline label auditing, parallax is the surface that still fires.

Don't train a better judge a policy can still fool. Add a judge it structurally can't — and read the disagreement as the signal.

Deterministic and reproducible; the benchmark is pre-registered and adversarial-symmetric; the frontier-VLM comparison is real, not hypothetical. Every number is reproducible from the named scripts in the vtl_spatial engine and its benchmark suite. Best fit: teams building success detection, reward labeling, or safety gating for embodied AI.