On real photographs, a deterministic two-view geometric check separates real 3D objects from flat printed decoys within a stated envelope — and catches ten flat decoys a frontier VLM certified as real objects present. It is a spatial GPS for embodied AI: it doesn't replace the VLM judge or the policy, it runs beside them, sub-second and un-gameable, and flags where appearance and geometry disagree — the candidate poisoned label. What follows is that result, and an honest map of what it catches, what it misses, and where a partner takes it next.
To scale training data for robot policies (VLAs), teams increasingly use a large VLM as an automated judge that scores whether an episode succeeded — shipped, e.g., in one production toolchain (VLM-as-judge episode scoring whose reward labels feed VLA fine-tuning). But the field documents how unreliable that judge is: on the RoboTwin 2.0 manipulation benchmark, a VLM success/failure observer scored precision 0.208 — of every five episodes it flags as successes, four are not — badly calibrated in both directions (RoboTwin 2.0). And robotics groups name the exact failure this work is about — VLA policies "fail silently: a collapsing rollout looks much like one making clean progress" (ValueFormer) — a live enough bottleneck that a 2026 cluster races to detect it (VLA-FAIL, SAFECAST). One study frames the mode in this paper's own terms — a Berkeley observability analysis of induced (not constructed) false successes finds "a peg that looks seated but is not" leaves almost no trace, corroborating that the silent-failure class occurs beyond printed decoys (observability study); notably, it also finds much of that signal is recoverable from proprioception — a point returned to in the ask.
Every one of those fixes is still learned — a better-trained judge is still a judge a policy can learn to fool. This work moves on a different axis: a deterministic, non-learnable geometric cross-check with no learned surface to attack. It doesn't try to win the judging contest; it grounds it. A spatial-GPS anchor that partners with the learned stack, where the two disagreeing is itself the signal. Independently, from image composition rather than robotics, it converges on the field's own stated bottleneck: silent failures poison the training labels.
Lineage · the composition origin — context, not the proof (the proof is Movement III)
Robots can see and robots can plan. What no standard pipeline instruments is whether the spatial intent of a natural-language instruction became a geometric fact. Semantic classifiers answer "does this look correct?" Point clouds handle collision geometry. Neither produces a scalar compliance gap when an instruction is satisfied perceptually but not geometrically.
The origin is an 8-dimensional composition kernel — a deterministic map from a frame to geometric field coordinates. As an aggregate fingerprint it carries real signal on the scenes it was built for: across compositional images, void ratio, packing, mass intensity and peripheral pull vary substantially and separate one arrangement from another (no single dimension is the story — the pattern across the vector, and the comparison between scenes, is). But it is worth being exact about which of it does the robotic work, because the robot task asks a narrower question. Real-vs-flat is a depth question — and a 2D composition metric cannot see depth at all; moreover, an isolated object on an empty table has no compositional arrangement for those dimensions to register, so on these tabletop scenes they sit near-constant across real and flat alike. So the robotic verdict rests on displacement (Δx, Δy) — did the mass land where instructed — plus a surface the kernel does not compute at all: an independent two-view parallax depth read. The honest description of the robotic instrument is displacement plus a depth cross-check; the kernel's other dimensions are live on the composition problem they came from, not dead weight — they simply do not answer the depth question the robot task poses.
K(I) = [ Δx, Δy, rv, ρr, μ, xp, θ, ds ] — a composition fingerprint
The eight are hand-designed features, not learned; fidelity as a proxy for "the object went where instructed" is measured here on constructed inputs, with validation on real rollouts still open. Stated positively: this is a spatial field, not a displacement gauge — its full multi-dimensional range is demonstrated on compositional scenes (Fig. 2 — void, peripheral pull, and packing all carry signal). The single-object robot demonstration simply does not exercise it, and that range applies naturally to scene-level robot verification — did the workspace get cleared, consolidated, aligned — the instrument's growth path, demonstrated on compositions, not yet on robots.
The kernel was developed on a different problem — spatial behavior in text-to-image generation. Across 6,000+ generated images on 10 platforms, spatial prompt intensity explains 0.0–0.1% of compositional-displacement variance; if instruction perfectly drove placement, the slope would be +1.000, and it isn't close. It is a real, independent finding about generative-model controllability, and the failure it names — a model satisfying a spatial instruction perceptually without executing it geometrically — is what pointed toward the robotics work.
The mechanism is a constraint substitution: models satisfy a spatial instruction by swapping in a perceptual strategy — clustering, dispersion, symmetry, partial placement — none of which guarantees the mass actually moved. In four hand-audited real scenes, humans said success, Gemini said success, and the kernel said failure. The kernel makes that gap a number.
Additional work · what was wrong, what we built, what we proved
The package shipped as a clean CLI that claimed to answer "did the robot put the object where the task required." Reading the source showed it did not. Fixing that — honestly, adversarially, and against a real frontier VLM — is the body of this work.
The kernel's "mass" was the 85–97th percentile of whole-frame gradient energy: table edges, shadows, clutter, the robot arm. So Δx/Δy was the centroid of scene texture, not the manipuland. In a clean single-object scene the two coincide; in any cluttered real scene the verifier fails silently and confidently — the worst failure mode there is. On a synthetic scene with known ground truth, the whole-frame centroid landed 0.559 (normalized) away from the true object, dragged toward background clutter.
Object isolation runs ahead of the kernel, in order of trust: an explicit segmenter mask › depth-plane removal › largest connected component. Perception may be learned; the verdict stays deterministic. This is the architecture the whole program implies — learned eyes, un-learned judge. Depth-plane isolation recovered the true centroid from 0.559 error to 0.000.
A single deterministic check is brittle; a single learned judge can be fooled by the very policy it grades. The fix reframes success itself: it is the agreement of N genuinely independent reads of one outcome, and their disagreement — Ω, the mean pairwise distance between the reads' object centroids — is the honest confidence.
Ω → 0: the surfaces agree — coherent, trustworthy. Ω large: the observables disagree — the outcome is contested, not passed. Trust no single read.
— the same coordinate the Parallax program uses in seismology and pathology, transposed to embodied verification
Independence is load-bearing: N reads that are all transforms of the same RGB surface measure transform diversity, not physical agreement. Each read declares its observable surface — rgb, depth, proprio, time — the audit counts distinct surfaces, and a single-surface verdict is capped at 0.5 confidence. Under a decoy attack (a flat printed object that satisfies appearance but has no geometry), Ω spikes to 0.81 and confidence collapses to 0.01: correctly contested, where a lone VLM judge is fooled.
We ran the parallax detector against two comparators on a labeled suite: an appearance-only proxy (RGB, geometry-blind) and the real Gemini Robotics-ER 1.6, given the same single RGB frame and a neutral judge prompt. The result is the headline.
The frontier embodied-reasoning VLM lands on the exact same score and the exact same error as the geometry-blind proxy: both false-pass every decoy. That dual result is the point — it confirms both that the proxy was a faithful stand-in, and that the blind spot is not a weakness of one model but of the single-RGB-frame setup itself.
"VERDICT: PASS — The dark block is located entirely within the bottom-left quadrant."
— Gemini Robotics-ER 1.6, on a flat printed decoy. It reasons correctly about the pixels; the pixels do not contain the fact that no object was placed.
The first suite only contained classes that break RGB judges — a cherry-pick. We corrected it: added classes designed to break our own depth surface, scaled to 1,200 scenes (120/class) with bootstrap 95% CIs, and pre-registered every prediction — including predicted failures of our method — before running.
| Class | Ground truth | Appearance (RGB) | Parallax (RGB+depth) | Who is right |
|---|---|---|---|---|
| clean_success | PASS | 100% | 100% | tie |
| wrong_place | FAIL | 100% | 100% | tie |
| clutter_success | PASS | 100% | 100% | tie |
| occluded_success | PASS | 100% | 100% | tie |
| empty | FAIL | 100% | 100% | tie |
| depth_noise_success | PASS | 100% | 100% | tie † |
| decoy (flat print) | FAIL | 0% | 100% | parallax |
| transparent_success (glass) | PASS | 100% | 0% | appearance |
| depth_dropout_success | PASS | 100% | 0% | appearance |
| raised_decoy (print on box) | FAIL | 0% | 0% | neither |
No judge dominates, and that is the honest conclusion. RGB wins where depth has no signal (transparent, dropout). Parallax wins on the flat decoy. And a printed photo mounted on a low box — raised_decoy — fools both surfaces at once. The safety metric that actually matters is the false pass: certifying a failed outcome as success poisons a training set and green-lights a bad action.
Read these synthetic per-class numbers as an upper bound. They come from a noiseless generator, where the flat-vs-raised separation is cleaner than any real sensor delivers — the same caveat the observability study makes about its own simulated separability. The load-bearing evidence is not this table; it is the real-pixel margin in Movement III (flat prints ~5–10 px vs real objects ~26–67 px, 86/89 in-envelope), measured on photographs with real noise.
Parallax eliminates the flat-decoy false pass entirely and halves the dangerous error — but does not reach zero. Read the two numbers precisely: it removes the camera-closable half (120 flat-decoy false passes → 0), and the entire residual (120) is raised_decoy — the one failure mode cameras cannot close, because both camera surfaces are fooled at once. So the residual is not a loose end; it is the exact size of what a third, non-camera surface must catch. Two independent surfaces are a floor, not a ceiling — and the 120 you can't close with cameras is precisely why the robotics partner is needed.
Two independent surfaces are a floor, not a ceiling — and the raised_decoy residual points at exactly what a third one has to be: a channel the cameras don't share. That is where a solo synthetic study becomes a collaboration, and it is taken up as the next step, not claimed as a result (below).
This verifier is the embodied-verification stage of a broader deterministic-measurement program — parallaxmetrology.com — but here it stands only on its own robotics evidence, above.
The synthetic proof, earned on real photographs — with the boundary stated
The capture happened. Real objects, photographed on an ordinary phone from two ~10 cm-offset positions — no depth sensor; the second view is the geometry. Run across 19 object types (real mugs, a paper cup, a Scotch-tape box and roll, and a range of printed decoys), the synthetic mechanism holds on real pixels.
Within a stated operating envelope, the instrument separates real 3D objects from flat prints with a wide margin: flat prints register ~5–10 px of relative parallax, real objects ~26–67 px — 86/89 in-envelope. The envelope is named, not hidden: the instrument measures height above the table plane, so transparent objects (a camera sees through them) and objects too short to parallax (a tape roll, a box lying flat) are out of it. A narrow instrument with a stated boundary, rather than a vague one. On transparency specifically, one coherent reading resolves what look like three: the instrument's intended behavior is to exclude glass as out-of-envelope; the varied empirical reads across tests (an anomalous single-frame raise, a synthetic false-fail, real detect-failures) are simply the noisy edge of that boundary, not conflicting results — a transparent object gives the depth surface nothing stable to measure, by design.
On a curated set of the strong flat decoys plus real controls, one RGB frame each went to Gemini Robotics-ER (neutral, object-named prompt) alongside the instrument. Gemini was fooled by 10 of 18 flat prints — and the instrument read every one as flat. The cleanest, threshold-robust cases are the flattened-box fake-3d (parallax 1.9–5.1 px, Gemini fooled 3/3) and the flat tape-box print (5.3–12.3 px, 3/3): comfortable geometric margins, on prints a frontier VLM certified as real objects present.
"PASS — a flat green Scotch Magic Tape package resting on the table." · "PASS — a cup is clearly visible resting on the table within the bounds of the white paper."
— Gemini, on two flat prints the instrument caught. It names the flatness / the paper, and certifies the object present anyway.
Two distinct failures live here, and they should not be conflated. When the decoy is RGB-identical to a real placement (the synthetic case, by construction), no RGB-only judge can tell — that is an information limit, and the second observable is genuinely necessary. But on the real prints, Gemini named the flatness — "a flat green package," "within the bounds of the white paper" — and passed anyway. There the information was in the frame and the model extracted it; its success-verdict simply ignored its own perception. That is not an information limit, it is a judgment failure decoupled from perception — arguably the stronger finding, because you cannot fix it by adding information. In both, a flat thing cannot fake the parallax shift, so the geometric check does not make the mistake; but the reason it wins differs — necessity of a second observable in one case, deterministic independence from a fooled judgment in the other.
The obvious objection: train a learned reward model on enough real data and it catches the decoys too. Two distinct answers, kept apart:
The decoy tests presence. We also ran the fake-free test of precision — a real cup at a ladder of offsets (0.6 to 28 cm) from a target — precisely because it should favor a measuring instrument over a VLM. It did not, and that is reported plainly. Gemini localizes as precisely as the instrument measures (its gap, read from its own two points, tracks the measured gap to a mean of 0.55 cm), and its verdict is sharp — it passes the dead-on placement and correctly fails everything from 2.2 cm out, with no lenient passes. On geometry it can perceive, a frontier VLM is a competent spatial judge, and the instrument holds no advantage.
That agreement is not a surprise — spatial localization is something embodied systems already do well — and it is not a consolation prize either. It is the calibration half of a working detector. The product is "run this beside the learned judge and flag the disagreements as candidate poisoned labels," and that requires two things: agreement when the outcome is fine (a low false-alarm rate) and disagreement when it is wrong (a true catch). Placement supplies the first; the decoys supply the second.
The two-surface residual (raised_decoy) is not a dead end; it is a specification. Robustness scales with the number of genuinely independent observables, and the missing one is the channel a camera cannot spoof: proprioception — did the gripper actually close on real mass of the expected profile? Touch, not vision. In a controlled study it closes every remaining hole: a blind depth surface (glass, occlusion) no longer vetoes an object the gripper can feel, and a raised decoy fails because a box has the wrong grasp profile.
Adjacent work — conceding the commodity floor, and naming the unoccupied combination
Geometry as a second signal for embodied AI is not new. The parts of this that are common should be conceded fast, because the novelty is a specific combination, not "add geometry."
The commodity floor. Using 3D structure to help a VLA is occupied. VeriSpace is a 3D-aware verifier that re-ranks candidate actions before execution; multi-view geometric consistency — the parallax mechanism used here — is classical structure-from-motion, table stakes rather than a contribution. Neither is claimed. This work does a different job (audit a completed outcome for a poisoned label, not select an action) with a different kind of check (deterministic, not learned).
What is unoccupied is the combination: a check that is deterministic (no learned surface for a policy to exploit), independent of the policy (so it can catch the policy's own perceptual errors), applied after the fact as a label-quality filter, where the disagreement between it and the learned judge is the signal.
The aspiration · from a verifier to a spatial-intent compiler
The current kernel measures displacement from the frame center. But the center is the wrong origin for almost every spatial instruction. "Place it in the corner" means the corner is the zero point — compliance should be distance from the intent, not from the frame. The delta has to move with the task.
That points past a verifier toward a different architecture. A natural-language instruction does not just define a target state — it initializes the coordinate system. Before the kernels run, the instruction is parsed to establish three things:
The kernels then read the field within that instruction-defined space. The same physical scene produces different vectors under different instructions — which is correct: compliance is always relative to intent. Add the Z axis — depth as a first-class dimension, not an edge case — and you have genuine 3D coordinates tied to intent.
Spatial instructions compile into measurement spaces, and execution is verified within them. The kernel stops being an approximation engine and becomes a coordinate engine — one component of a spatial-intent compiler.
— the design note this work opens onto
That is a larger and more consequential idea than a success detector, and the verifier is its first working component: the piece that measures, deterministically, whether intent became geometry.
Where the claim is a theorem, where it is falsifiable, and the one result that would kill it
The central claim — robustness scales with the number of genuinely independent observables — has to be pinned down, because part of it is true by construction and part is empirical, and conflating them is what makes it feel slippery. Kept apart:
The concrete niche is a label-quality filter in a VLA data engine: run this deterministic geometric cross-check alongside the VLM judge already scoring your episodes, and flag disagreement for review. It is cheap (sub-second, no GPU), reproducible, and has no learned surface for a policy to game — so it catches the confident false pass a better-trained judge would still wave through. Not a replacement for the learned stack: the ground-truth anchor beside it.
Two honesties that belong in the ask, not a footnote. First, the disagreement signal is trustworthy only inside the height-above-plane envelope. On transparent or camera-occluded objects the instrument disagrees when it is wrong — there it should defer, not fire. A deployable filter runs only where its observable is in-envelope. Second, mind the evidence tiers. The camera-only cross-check that halves the dangerous error is built and proven on real pixels; the residual it does not close (a 3D decoy that parallaxes like a real object) is closed only by a third surface — proprioception — which is modeled here, not built. Independent work (ProbeAct; the Berkeley observability study) suggests the touch/proprioception channel does much of the real recovery — corroboration for the multi-surface architecture, and the reason the third surface is the collaboration, not a shipped feature. It also invites the sharp question: if proprioception recovers most of it, is the built camera surface redundant? The honest parry: parallax is not the highest-recall channel, it is the independent one, and it is distinct in two ways proprioception is not. It is non-contact and works from a recorded video alone — so it audits logged episodes for label quality where no proprioception was captured (the data-engine case); and it is an independent geometric read inside the vision modality, so it catches a perceptual lie the learned vision judge shares and touch would only reach after the robot has already acted on the bad label. Where proprioception is logged, the two overlap on contact failures — conceded; where it isn't, or where the goal is offline label auditing, parallax is the surface that still fires.
Don't train a better judge a policy can still fool. Add a judge it structurally can't — and read the disagreement as the signal.
Deterministic and reproducible; the benchmark is pre-registered and adversarial-symmetric; the frontier-VLM comparison is real, not hypothetical. Every number is reproducible from the named scripts in the vtl_spatial engine and its benchmark suite. Best fit: teams building success detection, reward labeling, or safety gating for embodied AI.