Research lab book · measurement before interpretation

Question. Build. Question again.

This is the working record of an instrument being made: questions become constructions; constructions are attacked; failures become boundaries. The computational phase closes with no demonstrated differentiated biomechanical capability. The only remaining product hypothesis requires a prospective physical experiment.

COMPUTATIONAL PHASE CLOSED29 LINKED QUESTIONSPHYSICAL CLAIM NOT EARNED 2 EVIDENCE BLOCKERS552 / 560 TRIALS
Real corpus
552
included trials · 28 participants
Measurement table
357
auditable output columns per trial
Implementation verification
102 / 102
automated tests passing
Known-answer AUC
1.000
persistent plurality vs collapse
Orthogonal perturbation ΔAUC
−0.001
nonlinear nuisance · linear −0.018 · both fail
Structured-loss simulation
WITHIN CORPUS
AUC .912 · FPR 7.4% · no physical claim

01 · The instrument is the dialectic

Can organizational measurements reveal something consequential about the biomechanical system that produced them, beyond the complete incumbent read?
1 · QUESTION
one consequential unknown, named before the coordinate
2 · BUILD
the smallest instrument capable of answering it
3 · ATTACK
independence · null · resolution · ablation · incumbent
4 · CONSEQUENCE
held · killed · blocked · reframed
What is now executable. config/research_program.yaml is not a roadmap paragraph. It is an audited graph of 29 questions. A build fails audit if it has no adversary, claims a nonexistent artifact, skips its operating envelope, or ends without a next question. Run it with ztoy-biomech research-audit.
Built / tested
20
mode competence now passes clean gate
Blocked honestly
2
independent evidence depth · synchronized EMG
Open frontier
5
route persistence · external/device transport · precedence · cross-domain transfer

The quiver is plural

InstrumentQuestion it answersIt must not be collapsed intoState
Evidence depthHow many independent families must agree before a demand site exists?correlation or a composite confidence scorebuilt / DVJ blocked
PersistenceWhat survives changing sensors, timing, scale, and assumptions?strength of evidencebuilt
DependencyWhat other structures rely on a node?causal importancebuilt
Dynamic Ω / σ(Ω)How much do independent reads disagree, and is disagreement stable or intermittent?distance from normalbuilt / provenance blocked
Anatomical rotationDo strong local anchors migrate while the whole landing remains globally unsettled?peak timing or joint excursionbuilt / testing
Resolution-floor morphologyWhen does a coordinate become true, and how does truth arrive with resolution?a one-time robustness checkbuilt / unstable on DVJ
Structure-null protocolWhich destruction removes organization while preserving marginals?label permutationbuilt / direction diagnostic
Reference basinHow far has this system moved from a declared memory, and along which axis?constraint Ωbuilt
Reference LaboratoryIs viable movement one basin or several, and is this observation known, ambiguous, or outside all of them?one universal normal or an outcome-trained classbuilt / DVJ one basin
Directional deviation fieldWhere, when, and by what anatomical route does the system leave its memory?a scalar abnormality scorebuilt / consequence open
Backward claim integrityCan the headline be reconstructed from every untouched prediction and fitted object?a passing test count or mean fold scorebuilt / corrected headline
Direct consequence laboratoryDoes the quiver reduce direct internal-load waveform residuals beyond the complete incumbent?condition classification or tissue-stress languageevaluator only / no qualifying data
Orthogonal incremental valueWhat survives after incumbent-predictable outcome and instrument structure are removed?ordinary augmented-model gainbuilt / DVJ gate fails
Registered shift abstentionDoes the instrument refuse and localize degradation better than calibrated conventional comparators?beating a classifier with no abstentionbuilt / competitive gate fails
Plate-native two-family readHow do independently acquired marker and plate organizations disagree over landing phase?anatomical Ω or a third witnesstested / reliability mixed
Frozen-device transportDoes a competent load estimator fail safely on a physically unseen device without adapting to it?internal CV or arbitrary-shift coverageevaluator only / no qualifying data
Acquisition decision laboratoryWhich corpus is fit for which scientific job after access, license, sensing, target, and device gates?one mythical best datasettested / staged path selected
Registered capture QACan structured hardware loss be detected and localized at a subject-held-out operating point?arbitrary shift detection or a commercial device claimsingle-corpus simulation only
Lag-free periodic pluralityDo IMU, EMG-envelope, and EEG-electrical families share recording-specific frequency organization?pointwise synchronization, neural coupling, or one scale-free scalarwithin-corpus relation / scalar gate fails
Frozen local-atlas transportDoes local relational morphology survive unseen people and an entirely held walking speed?external apparatus transport or biological consequenceinternal gate passes
External local-atlas transportDoes the minimal PhysioNet IMU–EMG atlas survive a new apparatus and protocol?locomotion competence, mechanism, load, or clinical valueexternal gate killed

02 · Live questions, not a terminal hypothesis

After conditioning on the complete conventional kinematic, kinetic, force, coordination, and symmetry baseline, do relational measurements of disagreement and anatomical redistribution add held-out information about biomechanical constraint?
IDHypothesisRequired observationCurrent state
H1Independent observable families form a measurable phase×anatomy plurality.Nonzero Ω; plurality persists across scale; source-overlap audit passes.synthetic only
H2Constraint is intermittency as well as magnitude.σ(Ω), rolling sigma, and burst phase distinguish planted intermittent structure.held on controls
H3Added structure improves an incumbent-union model.Positive held-out ΔAUC and ΔBrier; bootstrap excludes zero; subject-aware permutation p≤.05.not held
H4The result is not a proxy for sensor reuse or reference definition.Provenance, every-read ablation, and block-ablation gates all pass.not held on DVJ
H5Anatomical localization is repeatable enough to interpret.Condition-specific ICC/CV/SEM/MDC and stable phase/node attribution.mixed

Question frontier

IDQuestion now openedConstruction or next buildEvidence state
BQ-001Do independent physical observables support the same demand sites, and at what epistemic depth?evidence_depth.pyblocked until ≥3 independent synchronized families
BQ-002Can a landing have strong local anatomical anchors without settling on one global anchor?rotation.pybuilt; planted route controls pass
BQ-003At what phase and anatomical resolution do Ω coordinates become stable enough to be true?resolution_floor.pybuilt; DVJ floor is unstable
BQ-004Which null destroys anatomical organization without changing marginal demand?null_protocol.pybuilt; null direction diagnosed by arm
BQ-007Does the morphology of the resolution curve identify the kind of organization being measured?resolution_morphology.pytested; controls, segments, and whole-subject source null pass
BQ-006Is compensation a stable route through anatomy or merely unrelated peaks?route survival + rotationopen
BQ-008Does the instrument improve direct internal-load estimation rather than condition labels?consequence_lab.pyevaluator built; planted control only; no qualifying data
BQ-009Does anatomical rotation precede conventional performance loss or follow it?fatigue / longitudinal precedence testnew data needed
BQ-010Which structures are deep, persistent, and load-bearing—and which are merely visible?joint Evidence–Persistence–Dependency readopen synthesis
BQ-011Where else does a field collapse from many viable organizations to one consequential path?new-domain transfer charteropen
BQ-012Does viable movement occupy one normative center or multiple stable organizational basins?reference_lab.pybuilt; DVJ supports one basin and refuses a taxonomy
BQ-013When a trial leaves a viable basin, where, when, and toward what organization does it travel?deviation_field.pybuilt; planted controls pass, biological consequence open
BQ-014Can every headline be reconstructed backward through folds, memories, provenance, and raw contracts?evaluation.pybuilt; fold-mean headline corrected to pooled OOF
BQ-015Does internal-load improvement survive an external cohort and unseen device?external device/site transport testopen; required before physical-load differentiation
BQ-016Does the instrument retain information after removing every component predictable from the incumbent?orthogonal.pybuilt; planted residual passes, both DVJ gates fail
BQ-017Is independence a corpus limitation or a sensing-topology boundary?INDEPENDENCE_BOUNDARY.mdreframed; kinematics-only can never clear plurality
BQ-018Can refusal and localization beat identically calibrated conventional comparators under degradation?shift_lab.pybuilt; current competitive gate fails
BQ-019Can DVJ support an acquisition-independent two-family read without promoting it to anatomical evidence?source_native.pytwo-source descriptive read computed; reliability mixed; no independent-plurality claim
BQ-020Does a competent direct-load model transport frozen to an unseen device and defer high-error trials?transport_lab.pyevaluator and planted control only; no qualifying physical corpus
BQ-021Must one corpus satisfy plurality, direct load, device transport, and commercial rights simultaneously?acquisition.pytested; no—staged falsification dominates the single-corpus bet
BQ-022Can the audit detect and localize structured capture failure at a usable held-out operating point?capture_qa.pyregistered injected-fault result; no physical or transport evidence
BQ-023Does the released PhysioNet corpus preserve enough synchronization to test plurality without manufacturing agreement?physionet_stage1.pyrecording-level admitted; pointwise phase and force closed
BQ-024Does lag-free three-family periodic organization survive source, subject, clock, artifact, segment, and resolution attacks?physionet_plurality.pywithin-corpus separation measured; registered scalar resolution gate fails
BQ-025Does scalar and cohort aggregation hide stable local organizations that reach the same total through different mechanisms?local_plurality_atlas.pyinternal topology measured; aggregate hiding demonstrated; taxonomy refused
BQ-026Does the frozen local atlas survive simultaneous unseen-subject and held-speed transport beyond scalar and global-curve summaries?atlas_transport.pytested; internal gate passes, IMU–EMG strongest, external claim refused
BQ-027Does the frozen PhysioNet IMU–EMG atlas retain recording identity under an external ENABL3S apparatus and protocol?enabl3s_transport.pykilled; AUC .513, no hidden subject/route/frequency island, target-native AUC .493
BQ-028Does local IMU–EMG organization add held-subject locomotion-mode information beyond the matched full incumbent and fail safely under channel loss?enabl3s_raw_mode_lab.py + enabl3s_fault_lab.pytested; confidence refinement narrow, classification and fault gates fail
BQ-029Can physical capture faults be detected and localized after a complete unseen-subject/device/site freeze?stage3_capture.pyadmission and freeze package built; physical corpus not collected
Binding language. This instrument measures registered external demand patterns. It does not estimate tissue stress, internal joint contact force, injury risk, or causality without direct validation against those targets.

03 · The trail — what changed and what it asked next

The lab book preserves wrong turns because they define the current gates.

2026-08-19 · orientation
Started in the wrong generation. The duplicate Ztoy folder was found to be stale. Work was redirected to mature documents/russell/ZTOYBOX; existing instruments were reviewed read-only.
opportunity map
Proposed a product hypothesis: not another force or symmetry scalar, but audited relational information—where independent observations disagree, whether it persists, and whether it adds information beyond incumbent coordinates. The computational phase did not demonstrate that hypothesis.
pilot v0 · insufficient reuse
The first implementation was too thin. It used conceptual vocabulary—Ω, dependency, persistence—but not the mature 50-coordinate field machinery, multiscale collapse, dynamic σ(Ω), full read ablation, configuration locks, or validation storm. This criticism was accepted and became the expansion brief.
structural port
Ported phase×anatomy geometry, persistent regions, islands, spectra, multiscale curves, dynamic Ω/σ, bursts, lags, correlations, plurality, and every-read ablation. Added moving-block uncertainty to preserve local phase dependence.
biomechanical reconstruction
Replaced generic marker speed with bilateral landmark-derived laboratory sagittal angles and continuous relative phase. Added center-of-pressure side assignment and side-aware external moment-magnitude projections. These remain external-demand proxies.
corpus pass #1 · 551/560
One extra exclusion was traced to a missing heel marker—not a missing trial. The adapter now uses a declared ankle fallback and records fallback count. Final yield returned to 552/560; eight source files are genuinely absent.
falsification found an audit defect
Two observable families produced no valid leave-one-family-out test but were being labeled stable. The gate was changed: fewer than three fused families is now explicitly not ablation-auditable, never silently stable.
first validation turn found outcome encoding
REAL30-reference features initially drove apparent perturbation AUC from 0.524 to 0.862. Because the outcome was REAL30 versus VR, distance-to-REAL30 encodes the definition. The block was removed from the primary benchmark and retained only as a labeled diagnostic. The later fold-mean rerun was 0.558 honest versus 0.841 diagnostic; the backward audit below subsequently replaced both headline estimators with pooled OOF values.
pilot turn closed, program opened
The first thirty tests passed and the public DVJ corpus did not earn differentiation. That result was not the end. It asked what the scalar Ω was hiding, when its coordinates became valid, and how many evidence families actually supported each demand site.
question → build: evidence depth
Saltimbanques’ assumption-depth instrument was brought into biomechanics. Each phase×node site now records how many families support it, the depth ladder, origin family, and all-family dropout. On DVJ, the coordinate is repeatable but cannot be called independent evidence because the sources overlap.
question → build: anatomical rotation
The art-domain local-anchor/global-field rotation question became a biomechanical one: can every moment be locally anchored while the whole landing refuses one global anchor? The new route instrument records anchor migration, travel, reversals, coverage, entropy, and local-minus-global anchoring.
question → build: resolution floor
Materials asked “when does the coordinate become true?” A phase×anatomy resolution ladder now measures that floor and its curve morphology. On DVJ the floor size has poor repeatability, so it is a calibration warning, not a new feature to celebrate.
Pathology interrogation · reference defects become gates
The pathology studies contributed the frame–memory distinction, directional drift, and the finding that distributions can beat mean deviation. Their notebooks also exposed dangerous failures: empty references that still save, self-cohort memories, silent mean imputation, and ambiguous Ω naming. Each became an executable refusal in the biomechanical Reference Laboratory.
question → build: Reference Laboratory
Multiple covariance-aware viable basins, leave-one-subject-out envelopes, bootstrap center uncertainty, destination ambiguity, unknown-regime output, frozen-memory serialization, and separate frame/memory fingerprints are now implemented. Known-answer two-regime controls pass.
real memory refuses invented phenotypes
The 27-subject REAL30 memory supports one stable basin. Twenty-four of 28 subject-excluded discovery runs select one; four select two. That instability is not converted into a movement taxonomy. Under the corrected 95% finite-sample conformal rule, 2.0% are unknown and 2.5% exceed the observed calibration range.
nested memory validation
Every memory was rebuilt entirely inside each outer subject-training fold. Pooled OOF clinical AUC moves from 0.421 to 0.426 (Δ +0.005); the outcome-anchored perturbation diagnostic moves from 0.531 to 0.510. The memory adds no current differentiating evidence.
backward claim audit · fold means fail reconstruction
The published headline used the unweighted mean of fold AUCs. Recombining every untouched prediction produced different primary values: clinical full-instrument ΔAUC changed from −0.185 to −0.122; honest perturbation changed from +0.033 to +0.021. The conclusion remains no differentiation, but the estimator was wrong for the headline. Pooled OOF metrics are now primary, every row must appear once, and permutation nulls reuse fixed observed folds.
tail-resolution adversary
A 97.5% empirical cutoff with 27 subjects claimed precision the calibration cohort cannot supply. The minimum conformal tail is 1/(n+1). The memory now uses a resolvable 95% envelope; undersupported basins return unresolved_outside_calibration. Frozen schema v2 memories carry an integrity fingerprint and reject tampering.
question → build: direct consequence laboratory
The next instrument is now executable: nested subject-held-out direct-load waveforms, complete incumbent versus augmented, target-sensor provenance refusal, and RMSE/R² plus peak, impulse, and timing residuals. The planted control improves RMSE from 0.3080 to 0.0397 and refuses target-source contamination. No physical-load claim exists until a qualifying external/device corpus is run.
question → build: orthogonal incremental value
The ordinary +0.021 perturbation delta was challenged directly. Both the outcome and all 200 added coordinates were residualized against the 127-coordinate incumbent inside subject-held-out training folds. The delta becomes −0.018 with linear nuisances and −0.001 with nonlinear nuisances; Brier calibration worsens. The observed gains also fail dimension-matched noise controls. The planted residual-signal arm passes, so this is a measured failure of the current DVJ block—not an evaluator that cannot detect signal.
independence boundary · corpus question reframed
Kinematics-only sensing cannot clear independent plurality regardless of sample size. Force plates are physically independent at acquisition, but the present anatomical load field reuses marker-derived moment arms. Acquisition independence, registration dependence, and target dependence are now separate contracts. A third synchronized physical source is a design requirement, not a desirable extra.
question → build: registered shift abstention
All 106 available REAL30 repeats were attacked with phase jitter, shared marker dropout, node substitution, force gain, and plate loss. The audit envelope produced median shift AUC 0.550 versus 0.594 conventional and abstained on only 0.7% of corruptions. Its competitive gate fails. The perturbation classifier is not a useful foil: it is already wrong on 92.6% of clean REAL30 subjects, and its split-conformal wrapper defers 85.2% of clean cases.
failure split · instrument versus substrate
The shift result contains two distinct failures. The audit coordinates are insensitive to three of five registered attacks; that is an instrument problem. The clean perturbation model is already unusable; that is a target/corpus problem. Selective prediction cannot be judged as graceful degradation until a clean competence gate passes.
question → adversary → rebuild: strict two-witness
Different transducers were not enough: net GRF is coupled by Newton's equations to marker-derived COM acceleration. It was removed from the witness score. COP speed and unordered inter-plate sharing remain. Across 552 trials, median JS is 0.261–0.284 bits and JS ICC only 0.124–0.577. Information-nonredundant two-witness measurement survives; reliability and anatomical plurality do not.
question → build: frozen-device consequence + refusal
A whole-device protocol now excludes the target device from fitting, tuning, normalization, residual calibration, and OOD thresholds. It reports zero-shot waveform error, empirical coverage, selective risk, and error enrichment. Base-task competence is a binding pre-gate. The shifted-device control passes; no qualifying physical multi-device load corpus has been run.
acquisition prescription adversaried
Scientific direct-load truth and a sellable route are disjoint. PhysioNet remains Stage 1 plurality/reliability; CAMS-Knee remains a six-subject noncommercial scientific kill test; rights-clean multi-device capture QA is Stage 3. Neither Stage 1 nor Stage 2 tests unseen-device transport.
question → build → run: injected-fault operating point
The two encouraging fault rows were forced through nested subject-held-out threshold calibration. The audit linear model reaches AUC 0.912, 7.4% clean false alarms, 74.1% fault sensitivity, and 90.7% localization versus conventional linear AUC 0.691 and sensitivity 42.6%. This is a registered result on injected faults in one corpus, not physical capture QA, transport, or commercial validation.
2026-08-20 · PhysioNet admitted under a narrower contract
All 713 publisher checksums verify. The extracted corpus contains 177 EEG, 169 EMG, 170 IMU, and 173 force recording keys; 164 contain the IMU–EMG–EEG trio and 162 contain all four modalities. EEG CSV and EDF are counted once, five subject-summary rows are excluded, and DOB is flagged for minimization.
released-clock adversary · synchronized does not mean traceable
Only seven of 164 EEG files retain a visible trigger transition; EMG is a 16.05 Hz MAV envelope, not raw EMG; three S47 IMU exports are empty; four more recordings exceed the 0.5 s duration-agreement gate. The surviving state is 157 recordings / 54 subjects for recording- or cycle-aggregate three-family work. Pointwise phase plurality and force-target work remain closed.
question → build → attack: lag-free plurality
IMU, released EMG envelope, and common-average EEG were converted independently to normalized gait-band spectra; no lag was estimated. Across 156 recordings / 54 subjects, observed mean JS is 0.203 versus 0.265 after speed-matched source permutation (p=.002). Every sensor pair and the within-subject cross-speed attack pass. One corrupt EEG channel map is refused.
frozen resolution gate refuses the scalar
Segment, fixed ±0.5 s clock-trim, planted, and common-mode attacks pass, but median resolution sensitivity is 0.093 bits against the frozen 0.05 limit. The separation survives 2× and 4× coarsening, while magnitude falls from median 0.184 to 0.160 to 0.091 bits. The phenomenon is multiscale; a single calibrated magnitude is not earned.
failure → new coordinate: resolution morphology
The scalar was not retuned. Four nested rungs now measure the curve itself. Fine, hierarchical, and coarse controls order correctly; empirical curves are monotone; median segment curve-AUC range is 0.023 bits. Observed curve AUC is 0.091 versus 0.111 under source permutation, and the 49-subject block null also passes (both p=.002).
aggregate adversary · the scalar hides mechanism
Cohort-mean spectra retain only 21.8–36.4% of mean recording-level disagreement across four lower-band edges. There are 197 pairs within 0.002 scalar bits; 69 have local-shape L1 above 1.0. One global number can therefore be reached through sharply different frequency-family organizations.
question → build → attack: local plurality atlas
Thirty-four channels remain nested inside three physical families. Median topology contains six islands over 18.65 effective bins. Matched segment-shape L1 is 0.516 versus 0.943 for speed-matched donors; the 49-subject repeated-measures null is 0.512 versus 0.943 (both p=.002). The best cluster silhouette is only 0.182, so no phenotype taxonomy is manufactured.
question → freeze → transport → adversary
The atlas was tested with both subjects and one complete walking speed held out. It reaches AUC .929 versus scalar .701 and global curve .728; combined .934. The same-subject/different-speed attack reaches .970. Pairwise ablation finds the strongest internal read in IMU–EMG topology: .945 pooled and .986 hard-adversary AUC, so EEG is not required. External apparatus remained the next burden.
external apparatus transport · recording identity killed
Ten ENABL3S archives were verified against publisher byte counts and MD5 values. AB156 was quarantined after informing the adapter. On nine untouched subjects and 792 same-person/same-route pairs, the frozen atlas reaches AUC .513 [.494, .533] versus scalar .519. No subject, route, or frequency island survives; minimum max-|t| familywise p=.610. A target-native held-subject diagnostic also fails at AUC .493 [.478, .508]. A complete hostile 500 Hz clock rerun remains null (atlas .484). The internal result was an apparatus fingerprint, not a portable identity coordinate.
physiological substrate recovered
ENABL3S five-mode classification passes the clean gate on nine untouched subjects. Equal-subject IMU balanced accuracy is .868 with macro AUC .978 and worst-subject accuracy .742; conventional EMG raises accuracy to .904. The feature table's missing circuit identity blocks relational attachment, so raw-circuit reconstruction is now mandatory.
raw relational increment · aggregate split
Exact two-second crops restore circuit and row provenance for 1,497 windows. Against the matched full conventional union, the atlas moves balanced accuracy from .928 to .933, but the paired subject interval crosses zero [−.012, .022] and only two of nine subjects improve. Log loss improves from .236 to .191 in seven of nine subjects (p=.037), although one subject supplies 74.2% of positive magnitude. Confidence refinement is retained; broad discrimination is not.
registered raw faults · localization without detection moat
Five faults are injected before feature extraction and thresholds are nested inside training subjects. Audit AUC is .958 versus .951 conventional, clean FPR 2.2%, sensitivity 44.0%, and localization 93.1%. The atlas adds only .006 AUC and misses most lag and left/right-substitution faults. The gate fails; refusal helps gain and swap but harms or does not help the other mechanisms.
failure → physical freeze package
BQ-029 now fails closed before collection: seven physical fault families at two severities per device; raw, synchronization, and independent fault-log hashes; hardware/firmware/software and ethics/consent provenance; rights; clean pairing; subject/device/site separation. Two devices can ask only named A→B transport; three or more can rotate the held device. Admission can never award a performance or commercial claim.
current state · not a terminus
One hundred two tests pass; the measurement frame remains fee2924dbd60ca54, Reference Laboratory configuration is a34d559cae069fa3, and frozen library integrity is 32697b2a9102d84b. The 29-question program has eleven built instruments, nine tested constructions, one killed external target, one reframed boundary, two blockers, and five open questions.

04 · Biomechanical Reference Laboratory

Does viable movement occupy one normative center or multiple stable organizations—and when a trial leaves, where does it go?
FIXED FRAME
declared coordinates
fee2924dbd60ca54
INSERTED MEMORY
config a34d559cae069fa3
integrity 32697b2a9102d84b
DEVIATION
distance · direction · route · unknown
CONSEQUENCE
held-out external criterion
Usable reference subjects
27
physical 30 cm landings
Stable viable basins
1
no supported phenotype taxonomy
Conformal unknown
2.0%
95% cutoff · 2.5% outside observed range
Center uncertainty
0.284
mean bootstrap standardized units
Evaluation scopeBaseline AUCWith memoryΔAUCInterpretation
Clinical · memory fitted inside outer subject folds0.4210.426+0.005pooled OOF; internal test only
Physical vs VR · same leakage control0.5310.510−0.021pooled OOF diagnostic; REAL30 anchors outcome
What the memory refuses. Subject self-reference; fewer than six subjects per basin; unstable or structure-null-equivalent partitions; any missing predeclared axis; candidate-basin mean imputation; corrupted frozen memory; covariance collapse; phase-grid, anatomical-node, observable-family, or configuration drift; false tail precision; and forced assignment outside an undersupported envelope.

Full contract and outputs: docs/REFERENCE_LAB.md and outputs/reference_lab_observation.json.

04b · Direct Consequence Laboratory

After the complete incumbent predicts a direct internal-load waveform, does the organizational quiver reduce untouched-subject residual error?
DIRECT TARGET
implant · tissue · validated simulation
PROVENANCE GATE
target sensor excluded from predictors
NESTED HOLDOUT
inner tuning · outer subjects
WAVEFORM RESIDUAL
RMSE · R² · peak · impulse · timing
Planted incumbent RMSE
0.3080
pooled untouched-subject waveform
With planted field read
0.0397
ΔRMSE +0.2683
Planted ΔR²
+0.630
localized residual recovered
Physical corpus
0
qualifying datasets evaluated
The line is explicit. The planted result validates evaluator behavior, not physiology. A physical-load claim remains false until a preregistered instrumented cohort and, where applicable, unseen device/site transport pass without target-driven recalibration.

Executable contract: src/ztoy_biomech/consequence_lab.py · protocol: docs/CONSEQUENCE_LAB.md.

04c · Registered Shift & Abstention Laboratory

Can the instrument win on trustworthy refusal under degradation when it does not win on classification accuracy?
REGISTERED ATTACK
jitter · shared dropout · substitution · gain · plate loss
UNTOUCHED SUBJECT
27 subject-held-out memories
STRONG COMPARATORS
conventional envelope · split-conformal classifier
RISK + REFUSAL
detection · abstention · localization
Conventional shift AUC
0.594
median across five attacks
Audit shift AUC
0.550
median · delta −0.003
Audit shift abstention
0.7%
too low to be protective
Localization
61.5%
mean registered-family accuracy
Registered attackConventional AUCAudit AUCAudit ΔAudit abstentionLocalization
Phase jitter 2%0.5010.498−0.0030%51.9%
Shared marker knee dropout0.5940.671+0.0770%66.7%
Shared ankle/knee substitution0.5030.501−0.0030%37.0%
Force gain 20%0.6050.550−0.0553.7%59.3%
Right plate loss0.7480.797+0.0490%92.6%
Headline refused. The trust-under-shift axis is not yet a demonstrated advantage. The perturbation classifier cannot rescue the story: it misclassifies 92.6% of clean REAL30 subjects, with 70.4% high-confidence errors. Its subject-calibrated split-conformal wrapper defers 85.2% of clean cases and produces no correct singleton decisions. That is a weak target, not safe selective prediction.
Do not bundle the failures. Shift blindness on phase jitter, node substitution, and force gain is an audit-sensitivity defect. The 92.6% clean error is a base-task/corpus defect. The first requires a sharper instrument; the second requires a learnable target. Neither is evidence that selective prediction is impossible.
Coverage language. Standard conformal coverage requires exchangeability. These registered corruptions are intentionally shifted from the clean calibration set. The results are empirical stress tests, not guaranteed coverage under arbitrary distribution shift. Weighted or nonexchangeable conformal methods require additional assumptions and data.

Protocol: docs/SHIFT_ABSTENTION.md · canonical outputs: outputs/shift_lab/.

04d · Registered capture-quality laboratory

Can the apparatus detect and localize structured capture failure at a usable operating point on an untouched subject?
REGISTERED FAULTS
knee-marker dropout · plate loss
NESTED SUBJECT CV
threshold learned inside training only
OPERATING POINT
≤10% training FPR
LOCALIZE
clean · marker · plate
Audit AUC
.912
linear · conventional .691
Clean FPR
7.4%
conventional 14.8%
Fault sensitivity
74.1%
conventional 42.6%
Localization
90.7%
conventional 77.8%
Feature familyModelAUCClean FPRSensitivityLocalization
Auditlinear0.9127.4%74.1%90.7%
Auditnonlinear0.9317.4%70.4%81.5%
Conventionallinear0.69114.8%42.6%77.8%
Conventionalnonlinear0.88814.8%66.7%85.2%
Registered within-corpus result. Thresholds are cross-fitted inside training subjects, and this injected-fault operating point improves false alarms, sensitivity, and localization over the matched linear comparator. It does not establish that the detector works on a physically degraded sensor.
Commercial claim refused. These are injected faults in one captured corpus. No physical fault recollection, unseen device, external site, or commercial-use prospective cohort has been tested. Stage 3 is designed to cross that boundary.

Protocol: docs/CAPTURE_QA.md · canonical outputs: outputs/capture_qa/.

04e · Plate-native strict two-witness laboratory

What remains when two reads must be independent in information—not merely measured by different transducers?
MARKER ONLY
mean motion · change · spread
PLATE NATIVE
COP speed · unordered sharing · change
GRF RESULTANT
coupled diagnostic only
HARD STOP
no joint · tissue · plurality claim
Trials
552
28 subjects · 8 source exclusions
Nonredundant witnesses
2
marker organization + COP/sharing
Median JS range
.261–.284
bits across five conditions
JS ICC range
.124–.577
poor to mixed repeatability
ConditionMedian JS bitsMedian phase transportMean phase correlationJS ICC(2,1)
REAL300.2840.0880.3530.322
VR00.2610.0800.3540.577
VR100.2770.0810.3340.124
VR300.2740.0840.3310.307
VR500.2810.0890.3590.402
What survived. COP motion and unordered inter-plate load sharing provide a plate-native witness not determined by whole-body COM acceleration. Amplitude, planted-shift, shared-source, information-coupling, and leave-one-read-out adversaries are explicit.
What did not. Different sensor physics alone was insufficient; net GRF was a coupled echo and has been removed from the score. The strict read still cannot localize demand, supply a third family, or clear evidence-depth. Median leave-one-read-out JS range is 0.112 bits (p95 0.366); no dose coordinate survives Holm.

Protocol: docs/NATIVE_TWO_FAMILY.md · canonical outputs: outputs/native_two_family/.

04f · Frozen-device consequence + refusal

On a learnable direct-load task, can the instrument remain useful—or know it is not—when the physical device changes?
SOURCE DEVICES
subject OOF competence
FREEZE
fit · tune · normalize · calibrate
UNSEEN DEVICE
zero-shot waveform
SELECTIVE RISK
defer high-error trials
Clean competence
PRE-GATE
no refusal claim if source R² fails
Target adaptation
NONE
including normalization and calibration
Shift guarantee
NO
empirical whole-device evidence only
Physical corpus
OPEN
known-structure control passes
The benchmark is now one question, not two roadmaps. Zero-shot internal-load error and audit-ranked refusal are scored on the same untouched device. A strong temporal model must be registered as the real incumbent; grouped ridge is only the executable reference control.

Protocol: docs/FROZEN_DEVICE_TRANSPORT.md · executable contract: src/ztoy_biomech/transport_lab.py.

04g · Acquisition prescription under attack

Should we buy the final corpus now—or force the instrument through cheaper, differently decisive corpora first?
CorpusSubjectsIndependent sourcesDirect loadCommercial pathDecision
DVJ-CAI282NoCC BYexhausted for current claims
PhysioNet multimodal gait594NoCC BY 4.0Stage 1 · aggregate plurality admitted
CAMS-Knee v1.164+implant forces + momentsscientific-onlyStage 2 · direct-load falsification
Knee Load Grand Challenge6 reported4implant contact forcefile terms unresolvedrefused pending terms + provenance
Prospective studyto register2 = named A→B; ≥3 = rotating holdoutphysical fault labels; direct load optionalrights securedStage 3 · commercial capture QA
STAGE 1 · SCIENCE
Does plurality exist and repeat?
STAGE 2 · LOAD
Separate noncommercial consequence test
STAGE 3 · QA
Independent rights-clean product route
Checksums
713/713
publisher SHA-256 manifest
Triggered trio
164
57 subjects before clock gates
Admitted
157
54 subjects · aggregate/cycle scope
Pointwise phase
CLOSED
released trigger traceability insufficient
Assertive decision. PhysioNet Stage 1 proceeds, but at recording/cycle level with clock-shift and EEG-artifact attacks. Use CAMS-Knee strictly as a noncommercial scientific direct-load kill test. The only commercially testable hypothesis is a rights-clean multi-device physical capture-QA study.
Released-data boundary. A common trigger existed during acquisition, but the processed exports do not retain equivalent trigger evidence. Fine phase alignment may not be reconstructed by maximizing cross-modal agreement. Force was manually co-started and its released timestamp/sync columns are non-informative; force-target analysis remains closed.

Admission: docs/PHYSIONET_ADMISSION.md · outputs: outputs/acquisition/physionet_admission.json and outputs/physionet/clock_report.json.

04h · PhysioNet lag-free plurality laboratory

Do independently acquired IMU, EMG-envelope, and EEG-electrical families share recording-specific periodic organization without manufacturing agreement through alignment?
SOURCE-NATIVE
IMU 60 Hz · EMG MAV 16.05 Hz · EEG 300 Hz
LAG-FREE
normalized Welch spectra · 0.5–6 Hz
THREE WITNESSES
JS · pairwise ablation · depth sweep
ATTACK
source · subject · segment · clock · artifact · scale
Measured recordings
156
54 subjects · one corrupt EEG map refused
Observed mean JS
.203
permuted .265 · Δ −.062 bits
Source permutation
p=.002
all three pairwise tests also pass
Frozen headline gate
FAIL
resolution Δ .093 > .05 bits
Nested resolution morphologyobserved recording match vs speed-matched EEG source permutation

The 45× rung is the registered one-bin collapse and is zero by construction. At the three informative rungs, observed-versus-source-permuted separation survives (each one-sided p=.002). The failed scalar-invariance gate is retained.

AdversaryObservationGate
Planted frequency + amplitudeShifted JS .875 vs aligned .003; amplitude delta 1.1×10−16pass
All pairwise speed-matched permutationsIMU–EMG Δ −.122; IMU–EEG −.086; EMG–EEG −.048 bits; each p=.002pass
Within-subject cross-speed EEGObserved .203 vs .331 bits; Δ −.128; p=.002pass
Three independent temporal segmentsMedian JS range .050 bits against ≤.10pass
Fixed ±0.5 s family trimsMedian maximum delta .002 bits against ≤.05pass
EEG common-mode comparatorMedian CAR minus common-mode JS −.021 bitspass
2× / 4× frequency coarseningMedian maximum delta .093 bits against ≤.05fail
What is now supported. Three separately acquired families contain recording-specific shared periodic organization beyond speed-matched source substitution and subject/session identity. Separation is present at every registered scale, while the scalar resolution gate fails.
What remains refused. A scale-invariant scalar; pointwise phase; brain or muscle mechanism; anatomy; tissue stress; internal load; and clinical value. Common-average referencing is an artifact adversary, not proof that gait-locked electrical structure is neural.
Resolution morphology earned. The failed gate changed the object. The four-rung curve passes known fine/hierarchical/coarse controls, has no nested-monotonicity violations, and is stable across three segments (median curve-AUC range .023 bits). Mean curve AUC is .091 observed versus .111 permuted (p=.002); a whole-subject donor permutation across 49 complete subjects gives the same −.020-bit separation (p=.002). Median half-collapse occurs at 9×.
Next burden. Transport the frozen curve to an external task or apparatus and ask whether its morphology carries consequence. The original single-resolution invariance gate remains failed; morphology does not rescue a scale-free scalar.

Protocol: docs/PHYSIONET_PLURALITY.md · canonical outputs: outputs/physionet_plurality/.

04i · Channel-aware local plurality atlas

When global totals look alike, are they governed by the same local organization—or does aggregation hide different frequency islands, family ownership, and channel structure?
3 FAMILIES
IMU · EMG envelope · EEG electrical
34 NESTED CHANNELS
substructure, never extra witnesses
LOCAL ATLAS
frequency density · ownership · islands
REFUSALS
band edge · donor segments · taxonomy
Median local islands
6
q75 support · 18.65 effective bins
Matched segment L1
.516
donor .943 · p=.002
Same-scalar pairs
197
69 have shape L1 >1
Taxonomy
REFUSED
best silhouette .182
How much disagreement survives cohort aggregation?family spectra renormalized symmetrically inside each band

The aggregate retains 21.8–36.4% of mean recording-level disagreement. Moving the lower edge does not remove the hiding effect.

Layer or adversaryObservationDecision
Same mass, planted topologyOne contiguous island versus three separated islands at identical aggregate masscontrol passes
Recording segmentsMean L1 .516 matched versus .943 speed-matched donors; p=.002stable local state
Whole-subject donor null49 complete subjects: .512 matched versus .943 donors; p=.002repeats respected
Lower edge .5–1.25 HzWithin-recording median L1 .494–.519 versus .948–.987 between recordingstopology persists
Channel heterogeneityMedian within-family JS: IMU .658 · EMG .215 · EEG .173deeper aggregate layer
Boundary peak111/156 recordings peak at .5 Hzcalibration warning
Unsupervised typesBest k=2–6 silhouette .182no phenotype classes
Internal measurement result. The global scalar and cohort mean hide local differences within this corpus. Local frequency-family topology is internally stable and distinguishes many recordings whose global JS is essentially identical; later external transport fails.
Mechanism still refused. Stable recording-state topology is not a subject phenotype, neural signature, anatomical route, internal load, or clinical value. Channel differences remain nested within their physical family.
Next step preserved and sharpened. Freeze the four-rung resolution curve and the local atlas, then transport both unchanged. Transporting only the aggregate would validate the hider rather than the instrument.

Protocol: docs/LOCAL_PLURALITY_ATLAS.md · canonical outputs: outputs/physionet_local_atlas/.

04j · Frozen local-atlas internal transport

If the aggregate is the hider, does the local organization still identify a recording when both the person and the walking speed are unseen?
TRAIN
other subjects · other speeds
FREEZE
scaling · coefficients · thresholds
TEST
unseen subject + held speed
ATTACK
same subject · other speed · channel pairs
Scalar AUC
.701
aggregate local JS mass
Global curve AUC
.728
four resolution rungs
Local atlas AUC
.929
held person + held speed
IMU–EMG atlas
.945
strongest pair · EEG excluded
What survives the double holdout?same-recording segment matching · pooled out-of-fold AUC

The combined model reaches .934 but does not beat the atlas alone; accumulation is not the result. IMU–EMG local morphology is the strongest read.

AttackObservationDecision
Zero target-speed fittingEach 0.5, 0.75, or 1.0 m/s test excludes that complete speed and every test subject from fittingpass
Held-speed floorCombined AUC .933, .970, .933 across the three speedspass
Scalar incumbentCombined ΔAUC .233; left-subject-block bootstrap 95% CI .195–.277local structure retained
Global-curve incumbentCombined ΔAUC .205; 95% CI .172–.243locality matters
Same subject, another speedAtlas AUC .970; IMU–EMG AUC .986not explained by person identity
Pairwise family ablationIMU–EMG .945; IMU–EEG .875; EMG–EEG .875EEG not required
External apparatusAll recordings originate in one PhysioNet apparatus and tasknot tested
Internal result only. Inside this apparatus, frequency-local IMU–EMG relational morphology retains recording-specific structure across unseen subjects and a held walking speed that scalar and global-curve summaries discard. The subsequent ENABL3S test shows this is not externally portable.
What is not earned. This can still reflect task-, session-, or apparatus-specific spectral consistency. It is not a muscle mechanism, neurological coupling, anatomical constraint, internal load, phenotype, or clinical value result.
Subsequent decision. The frozen IMU–EMG atlas was moved unchanged to ENABL3S and collapsed to chance. The internal result is retained as apparatus-specific recording discrimination, not transport.

Protocol: docs/ATLAS_TRANSPORT.md · canonical outputs: outputs/physionet_atlas_transport/ · contract bbd73c72ff2394e9.

04k · Frozen external local-atlas transport

Does the PhysioNet IMU–EMG atlas survive when the apparatus, protocol, participants, EMG representation, and recording structure all change?
SOURCE
PhysioNet only · AUC .909
FREEZE
channels · grid · model
QUARANTINE
AB156 adapter development
TEST
9 untouched ENABL3S subjects
Verified archives
10 / 10
publisher bytes + MD5
Evaluation pairs
792
same person · same route parity
Frozen atlas AUC
.513
95% CI .494–.533
Target-native AUC
.493
held-subject diagnostic
AttackObservationDecision
Development leakageAB156 informed the 2.5-second crop contract and is excluded from the gatequarantined
Scalar comparatorScalar AUC .519 [.508, .530]; atlas-minus-scalar −.006 [−.034, .020]no local increment
Subject decompositionAtlas AUCs .451–.569 across nine untouched peopleno stable subgroup
Route mechanismEven .533; odd .489no repeatable route split
Frequency islands12 bin AUCs .487–.520; minimum max-|t| familywise p=.610aggregate not hiding a local win
Target-native substrate12-bin ENABL3S-trained model, held subjects: AUC .493 [.478, .508]recording identity unusable
Multirate clock ambiguityHostile 500 Hz rerun: atlas .484 [.459, .508], scalar .499, target-native .479null invariant to clock contract
BQ-027 is killed. The PhysioNet atlas is a strong internal recording fingerprint, but it does not transport to ENABL3S. The null is not a hidden-island cancellation, and a target-fitted held-subject diagnostic finds no usable recording-identity substrate either.
Dialectical consequence. Stop tuning identity. Use ENABL3S for its physiological labels: locomotion modes and transitions. The next build compares an IMU-only incumbent with local IMU–EMG organization under subject holdout, then attacks both with registered channel loss and confident-error/refusal tests.

Protocol: docs/ENABL3S_EXTERNAL_TRANSPORT.md · canonical outputs: outputs/enabl3s_atlas_transport/ · contract 9fa77dc21ea54719.

04l · ENABL3S physiological competence gate

Before asking whether an instrument refuses degradation gracefully, is there a clean held-subject task on which it is genuinely competent?
Evaluation subjects
9
AB156 development-only
Stable windows
15,410
five locomotion modes
IMU balanced accuracy
.868
equal-subject · worst .742
IMU + EMG
.904
macro AUC .992
RepresentationBalanced accuracyMacro AUCWorst subject
Published IMU features.868.978.742
IMU + conventional EMG.904.992.776
IMU + goniometer + EMG.965.999.899
Clean competence passes. Unlike DVJ perturbation and ENABL3S recording identity, this target supplies a working held-subject baseline from which degradation can be measured.
Backward refusal. The publisher feature table lacks circuit identity. It may establish subject-held competence, but it cannot receive a guessed raw-waveform atlas or support circuit-block uncertainty. Relational and channel-loss arms must be rebuilt from circuit CSVs with archive/member/row provenance.

Protocol: docs/ENABL3S_MODE_LAB.md · canonical outputs: outputs/enabl3s_mode_lab/.

04m · Circuit-traceable mode and raw-fault laboratory

When the aggregate moves, is it broad information, confidence refinement, or one subject and one fault mechanism carrying the result?
Exact raw windows
1,497
1,323 held-subject · two seconds
Matched accuracy Δ
+.005
95% CI −.012 to .022
Matched log-loss Δ
−.045
7/9 improve · p=.037
Fault AUC increment
+.006
audit .958 · conventional .951

Clean relational increment

RepresentationBalanced accuracyLog lossAdversarial reading
IMU only.920.257strongest sparse sensor read
IMU + EMG.892.303EMG harms this fixed model
IMU + EMG + atlas.913.248recovers damage; does not beat IMU
Full conventional.928.236matched primary incumbent
Full + atlas.933.191confidence improves; accuracy CI crosses zero
The aggregate was hiding a different mechanism. It was not hiding broad accuracy. Only two subjects improve on balanced accuracy. Log loss improves in seven, but 74.2% of positive magnitude comes from one subject. The atlas currently adjusts confidence more credibly than decisions.

Fault-by-fault detection: nonlinear audit arm

Registered raw faultFlag rateLocalizationInterpretation
Waist IMU dropout1.0001.000hard loss found
Right TA electrode dropout1.0001.000hard loss found
EMG bank gain ×2.638.967partial detection
Left/right IMU substitution.241.819mostly missed
EMG lag 100 ms.186.675mostly missed

BQ-028 decision: substrate yes; differentiator no

The competent five-mode task survives, and the relational read shows a bounded confidence-refinement signal. But it does not earn broad incremental discrimination, and the audit detector does not separate materially from the conventional detector. Strong localization is conditional on finding the fault. Further ENABL3S tuning would optimize known simulations; the evidentiary move is rights-clean physical faults. Two devices support one named A→B transfer; three or more permit rotating held-device tests.

Protocol: docs/ENABL3S_RAW_MODE_AND_FAULT_LAB.md · clean artifacts: outputs/enabl3s_raw_mode_lab/ · fault artifacts: outputs/enabl3s_fault_lab/ · contract 3b974e5dbcd79e44.

04n · Stage-3 physical capture freeze

What must be frozen before physical data exist so that a friendly dropout demonstration cannot masquerade as device-general capture QA?
Physical fault families
7
each device · ≥2 severities
Two devices
A→B
named transport only
Rotating holdout
≥3
still not device-general proof
Claims earned by admission
0
evaluation permission only
Admission surfaceExecutable requirementOverclaim prevented
Fault truthPhysical realization plus independently hashed fault logself-declared or feature-masked “physical” fault
Synchronizationevent ID, timestamp source, and hashed sync logoutcome-optimized alignment
Device identitymanufacturer, model, serial hash, hardware, firmware, softwaredifferent IDs for the same acquisition surface
Fault coveragedropout, intermittency, gain, substitution, offset, drift, attachment/occlusioneasy-dropout benchmark sold as general QA
Transport unitsubject-disjoint devices; site scope tracked separatelypooled trials or subjects counted as device replication
Rights and governancecommercial development/validation, derivatives, audit retention, ethics and consent versionsscientific access converted into a product asset
Safety boundary. Faults may not alter participant-facing control or safety-critical feedback. On assistive or robotic platforms, inject them into a shadow acquisition path unless a separate safety protocol explicitly authorizes otherwise.
Study-size boundary. Three subjects per device is only the executable minimum. The planning target begins at four platforms across at least three sites and 20 subject-disjoint participants per platform, then increases through clustered precision simulation until clean-refusal and fault-sensitivity confidence bounds can resolve their gates.
What is ready. The frozen YAML protocol, empty manifest template, checksum/provenance auditor, CLI, and known-answer attacks are built. What remains external is SOP rehearsal, precision simulation from rehearsal variance, ethics/rights execution, hardware/site selection, registration, and physical collection.

Protocol: docs/STAGE3_PHYSICAL_CAPTURE_PROTOCOL.md · freeze: config/stage3_capture_protocol.yaml · manifest: templates/STAGE3_CAPTURE_MANIFEST_TEMPLATE.csv · YAML fingerprint 6579c6b6bba41966 · executable contract 0cdccb54993650cb.

04o · Computational phase synthesis — paper wrap

After every backward audit and external adversary, what actually survives—and what experiment can still change the claim?
Tier earned
1
deterministic measurement
Domain result
NO WIN
no differentiated biomechanical capability
Remaining hypothesis
CAPTURE QA
physical performance untested
Next evidence unit
DEVICE
not pooled trials

Headline result

No differentiated biomechanical capability was demonstrated. The software deterministically measures, reconstructs, plants, attacks, and refuses. DVJ incremental discrimination, dose, evidence depth, and broad shift claims fail or remain blocked. PhysioNet local organization survives internal transport but dies on the external ENABL3S apparatus. ENABL3S is a competent substrate, yet relational accuracy and broad fault advantages fail; confidence refinement is narrow and magnitude-concentrated. Physical capture QA is the only remaining product hypothesis, and it has not been tested physically.

FindingStatusWhat may be said
Measurement, provenance, backward reconstruction, planted controlssurvivesTier-1 apparatus is functioning and auditable
ENABL3S log-loss refinementbounded0.045 mean improvement; interval 0.004–0.114; 7/9 subjects, but 74.2% of positive magnitude from one subject
Conditional fault localizationbounded93.1% among faults the detector finds; not unconditional detection
DVJ structured-loss simulationwithin corpusAUC .912 vs .691; clean FPR 7.4% vs 14.8%; sensitivity 74.1% vs 42.6%; no physical QA established
DVJ clinical/perturbation increment, dose, stable coordinatesnot earnedNo classifier, mechanism, or dose claim
External atlas transportkilled herePhysioNet-to-ENABL3S AUC .513; internal identity did not transport
Broad ENABL3S fault advantagefailsAUC advantage .006 and sensitivity 44.0%; gain, lag, and substitution remain weak
Physical/device-general/clinical/commercial performanceuntestedNo statement is authorized

The physical experiment, in one page

LayerRequirement before the claim can move
GovernanceEthics, versioned consent, commercial development/validation and derivative rights, retention, sponsor/data-controller, site and hardware agreements
TopologyTwo platforms for one named A→B test; preferably four across ≥3 sites for an initial rotating architecture; more device units for device-general inference
ParticipantsSubject-disjoint device cohorts, representative capture difficulty, ≥2 clean repeats per subject-task; final n set by clustered rehearsal precision, not “20” by convention
TasksStandardized walking, sit/stand, step-up/down, and a safety-approved higher-transient task with scripted events and recovery
Fault matrixAll seven physical families on every device, ≥2 calibrated severities: dropout, intermittent loss, gain, substitution, offset, drift, attachment/occlusion
TruthIndependent reference clock/packet/calibration logger, second-operator record, randomization, blinding, and SHA-256 raw/fault/sync evidence
FreezeFeatures, model, normalization, thresholds, taxonomy, tasks, severity, estimands, multiplicity, missingness, sample size, software and target-data embargo
Primary gatesClean false refusal ≤10%; sensitivity ≥70%; localization ≥70%; audit AUC advantage ≥.05; task competence first; worst fault and worst device headline
InferenceSubject clusters for named transport; devices as units for multi-device claims; site separate; no pooled-trial pseudo-replication
Kill ruleIf the frozen audit ties the complete incumbent, remains blind to informational faults, or wins only on one easy dropout/device/operator, close the current product wedge
Readiness verdict. The study is computationally prepared but not collection-ready. The immediate external milestone is governance plus a bench rehearsal on one development platform. Participant rehearsal then supplies cluster variance and score correlation for final precision planning. Only after that freeze should untouched target devices be opened.
Operational handoff built. In addition to the executable capture manifest, three empty prospective templates now bind the platform/site register, physical fault and safety SOP matrix, and development-rehearsal metrics needed for clustered precision. They do not invent sample size; they ensure the rehearsal captures what the final calculation requires.
Device-count boundary. Four platforms are an operational starting architecture, not a magic sample size. They enable rotating held-device attacks but may still be too few for a stable device-population interval. If device-level precision is inadequate, acquire more devices or narrow the wording to named transports.

Full synthesis: docs/COMPUTATIONAL_PHASE_SYNTHESIS.md · physical design and resources: docs/PHYSICAL_EXPERIMENT_REQUIREMENTS.md.

05 · What the instrument quiver reads

PHASE × ANATOMY FIELD footanklekneehip pelvislumbartrunk phasecontact demand path / centroid wander
L0Registration201 phase points × seven anatomical nodes; disclosed 33-point anatomical regularization for topology.
L1ConventionalForce, impulse, loading rate, bilateral angles, velocity, asymmetry, relative phase, waveform summaries.
L2RelationalΩ mean/sigma, burst, plurality, lag, correlation, source overlap, every-family ablation.
L3StructuralGeometry, void/cohesion, persistent regions, islands, contour, gradient, spectral bands.
L4ReferenceWithin-person REAL30 envelope, node/phase exceedance, L1 and Jensen–Shannon distance.
L5ValidationReliability, dose, block ablation, held-out models, bootstrap, permutations.
L6EpistemicEvidence depth, family origin, depth ladder, source overlap, all-read loss.
L7DialecticalAnatomical rotation, structure nulls, resolution-floor curve, next-question audit.
L8Frame + MemorySupported viable basins, covariance-aware direction, uncertainty, destination, unknown regime, phase×anatomy deviation route.
L9Orthogonal valueCross-fitted outcome/instrument residuals, nuisance-family agreement, duplicate and matched-noise controls.
Markers
bilateral landmarks
Joint motion field
|angular velocity|
Dynamic Ω
centroid · spread · entropy
Audit gates
plurality · provenance · ablation
Force plates + COP
side assignment
External load field
|r × F| / body scale
Field kernel
34 coords per field + aggregate
Held-out tests
incumbent vs augmented
Why the real Ω gate fails. External-load moment arms use the same marker coordinates as joint kinematics. This source overlap means the two fields are not independent. There are also only two fused families, so a complete leave-one-family-out audit is impossible.

06 · What came from existing ZTOYBOX instruments

SourceMature mechanism reusedBiomechanical adaptation
Material/vtl_materials_kernel.pyWeighted field geometry, anisotropy, void/cohesion, persistent regions, islands, contour/gradient, spectra.Applied to phase×anatomical demand fields; 34 coordinates per family and aggregate.
Saltimbanques/.../persistence.pyPlurality, multiscale survival, collapse, separate Ω and σ(Ω).Phase/anatomy scale ladder with “never collapses” retained as an explicit state.
persistence_regions.pyh-maxima region survival and monotonic count audit.Persistent demand regions, anchor emergence, collapse depth, count curves.
parallax_portrait/omega.pyRobust diagonal basin, axis attribution, configuration refusal.Subject-mean median/MAD basin plus hard config-fingerprint mismatch error.
Markets/starter_independence_audit.pyCorrelations, Ω curve, rolling sigma, lag, all-read sensitivity.Centroid/spread/entropy reads at every phase; pairwise and aggregate curves.
seismology/bootstrap_stage1.pyBlock bootstrap respecting autocorrelation.Lag-one phase autocorrelation sets bounded block length for Ω confidence intervals.
Material validation briefsRepeat aggregation, incumbent union, nonlinear residual, grouped validation.Subject×condition aggregation; 127 incumbent vs 200 instrument coordinates; complete subjects held out.
Pathology frame–memory and drift studiesInvariant measurement frame, replaceable contextual memory, directional failure, distribution over mean.Multiple supported basins, separate fingerprints, unknown-regime refusal, directional phase×anatomy routes, and memory fitting inside outer folds.
Pathology delta logic + backward auditNormative expectation, directional residual, and reconstruction from untouched predictions.Cross-fitted removal of incumbent-predictable outcome and instrument structure before an added-information claim.

Full implementation mapping: docs/ZTOYBOX_LINEAGE.md.

07 · Corpus, provenance, and accounting

Trial accounting by conditionincluded n=552

Expected: 112 trials per condition. REAL30 has six missing source trials; VR0 and VR50 have one each. Every manifest row is retained in the accounting table.

Registered archive
Zenodo 18503500
CC BY 4.0
MD5 a3b7e2c0ed48d727f434dfa8bfae380d
3,097 files · 3.95 GB expanded

Cohort
16 / 12
control / chronic ankle instability
Integrity itemObservationAction
Missing source files8 `.mot`/`.trc` paths absentExcluded by explicit reason; not imputed.
Marker alternatives30 trials use ≥1 declared fallback; 6 use twoFallback count included in measurement table.
EMGPresent for a subset, but no sample rate, time, or sync eventNot phase-fused; blocker stated in summary.
Configuration driftSettings could otherwise change under same metric namesEvery row stamped fee2924dbd60ca54; basin mismatch raises.
Runtime101.25 s for manifest; 0.181 s/trialRecorded after adding evidence, rotation, and resolution instruments.

08 · Known-answer falsification arms

Ω magnitude and temporal dispersion100 repetitions / arm
Ω meanσ(Ω)

The dependent-source arm is intentionally identical in geometry to persistent plurality. It is rejected only by provenance—demonstrating why structural signal alone is insufficient.

Plurality vs collapse
AUC 1.0
complete separation
Dependent source refusal
100%
despite high, persistent Ω
Burst sigma success
100%
burst σ exceeded collapse
Arm definitions

Collapsed: independent reads share one anatomical trajectory. Persistent plurality: nine independent redundant reads occupy three separated anatomical trajectories and survive scale/ablation. Dependent plurality: identical geometry but all reads declare a shared source. Intermittent burst: one read departs only near phase 0.58.

09 · Industry comparison: did it add information?

Subject-held-out AUC: incumbent versus augmentedtechnical repeats aggregated first
127-feature incumbentaugmented primaryoutcome-anchored diagnostic

Pooled predictions from every untouched row are primary; fold means are diagnostic. Complete subjects are held out. For perturbation, REAL30-reference coordinates are excluded from the primary model because they encode the outcome definition. The violet diagnostic shows the misleading result if they are included.

Clinical block additionsΔAUC vs incumbent
Perturbation block additionsdiagnostic reference shown
Outcome / modelBaseline AUCAugmented AUCΔAUC95% bootstrap CIPermutation
CAI · logistic0.4210.298−0.122[−0.272, 0.019]p=.925
CAI · nonlinear0.5490.517−0.032[−0.153, 0.073]
Physical vs VR · logistic, honest primary0.5310.552+0.021[−0.073, 0.110]p=.323
Physical vs VR · nonlinear, honest primary0.5920.566−0.027[−0.139, 0.080]
Physical vs VR · reference-inclusive diagnostic0.5310.821+0.290[0.141, 0.428]not eligible

Orthogonal incremental-value attack

The ordinary comparison asks whether refitting with a wide new block helps. The stronger test removes everything in the outcome and added block predictable from the incumbent, learns only from residuals inside training subjects, and tests untouched subjects once.

Outcome / nuisanceBaseline AUCOrthogonal AUCΔAUCΔBrierClipped
CAI · linear0.4210.381−0.040−0.08465.5%
CAI · nonlinear0.5490.381−0.167−0.12828.8%
Physical vs VR · linear0.5310.513−0.018−0.14438.1%
Physical vs VR · nonlinear0.5920.592−0.001−0.00824.5%
The current gate is not cleared. The observed linear deltas do not beat matched-noise maxima (+0.133 clinical, +0.060 perturbation); calibration worsens; nuisance families do not yield a positive residual gain. Exact-duplicate incumbent deltas stay inside the declared ±0.05 sanity band. The planted control does pass: AUC 0.635 → 0.989, Δ +0.353, bootstrap lower bound +0.207, maximum matched-noise delta +0.001.

Decision at this turn: no residual information demonstrated on this target and corpus

The current DVJ quiver measures additional organizational coordinates, but this 28-subject study has not demonstrated outcome information orthogonal to the incumbent. The result does not establish absence of information in another target, sensing topology, or adequately powered cohort. The ordinary +0.021 perturbation result is compatible with redundant re-expression, finite-sample refitting opportunity, or unresolved small-sample variation. The program now carries a harder admission gate and moves toward direct internal-load consequence, genuinely independent synchronized evidence, and external device/site transport.

10 · Reliability and dose behavior

ICC(2,1) by conditionselected focal coordinates
<.40 poor.40–.59.60–.74≥.75

Reliability is a property of each coordinate under each condition, not of the instrument as a whole. The canonical table contains 100 condition×metric rows with CV, SEM, and MDC95 in addition to ICC.

Ordered VR-height responsemean within-subject Spearman ρ and 95% bootstrap CI
Multiplicity result.

Peak force and loading rate show unadjusted negative trends, but no focal coordinate survives Holm correction across the declared family. Dynamic Ω mean and sigma show no ordered dose behavior.

The physical 30 cm condition is not part of the ordered VR-height series.

11 · Claim ladder and present boundary

TierWhat may be saidEvidence requiredStatus
1 · MeasurementSoftware deterministically computes the named registered quantities.Tests, fingerprints, accounting, known-answer arms.earned
2 · AssociationA coordinate differs by condition/group in this corpus.Grouped estimates, uncertainty, multiplicity.mixed/exploratory
3 · Incremental valueAdded coordinates retain information beyond the complete incumbent union.Positive orthogonal AUC and Brier deltas across nuisance families, bootstrap, fixed-fold null, duplicate and matched-noise controls.DVJ orthogonal gate fails
4 · MechanismA stable compensatory path is localized.Independent synchronized families and stable phase/anatomy attribution.not earned
5 · Physical loadInternal force or tissue loading is better estimated.Independent direct target, target-source exclusion, nested subject holdout, peak/impulse/timing residuals, and external/device transport.evaluator built; data not tested
6 · ClinicalPatient decisions or harm prediction improve.Prospective external clinical validation.not tested

12 · Where we are going

Stage 1d closed · physiological target, no registered-fault moat.
ENABL3S supplies a competent five-mode substrate. The atlas refines confidence but does not add broad accuracy, and its registered-fault detector essentially ties the conventional union. Preserve the result; stop tuning these simulations.
RESULT: bounded signal + failed gates
Stage 2 · CAMS-Knee scientific kill test.
Request access under its noncommercial terms. Freeze the construction before opening implant targets; test direct-load residuals across six subjects and multiple activities. Treat low power as a bound, not transport evidence.
GATE: direct-load residual information
Register the strongest temporal incumbent.
Use raw-waveform LSTM/CNN/Transformer predictions or an immutable adapter under the exact whole-device freeze. Ridge remains a control, never the adversary we claim to beat.
GATE: non-strawman incumbent
Stage 3a · governance and bench rehearsal.
Name the sponsor, rights owner, platforms, sites, safety lead, and independent truth system. Realize all seven faults on a shadow/bench path with calibrated physical severities, start/end clock truth, recovery checks, and hashed evidence before participant recruitment.
GATE: safe independently logged faults
Stage 3b · quarantined participant rehearsal and precision.
Use development hardware only. Estimate clean false refusal, task competence, every-fault sensitivity/localization, device/subject variance, paired detector correlation, attrition, and operating burden. Freeze final n and platform count from confidence-bound resolution; 20 per platform remains only a seed.
GATE: resolvable clustered intervals
Stage 3c · freeze once; open untouched hardware once.
Lock tasks, fault SOP, severities, models, memories, thresholds, multiplicity, missingness, target embargo, and fingerprints. Collect subject-disjoint target devices/sites and report every device, task, severity, and fault. A target-phase repair becomes a new prospective version and cannot erase the primary failure.
GATE: physical + named/rotating transport
Frozen external normative memory.
Build a larger task- and apparatus-declared memory, freeze it before evaluation, and test whether multiple viable basins survive resampling and external subjects. Unknown observations must remain unknown rather than being forced into the nearest basin.
GATE: stable regimes + external calibration
Physical degradation recollection.
Keep the registered simulations for diagnosis, but recollect timing, dropout, substitution, plate loss, and device-change faults physically. Improve the three shift-blind audit surfaces without tuning against the held-out device.
GATE: sensitivity + clean false-deferral
Prospective clinical question—only after tiers 1–5.
Define a decision and outcome first; evaluate incremental utility, calibration, subgroup stability, and consequences. Do not retrofit “injury risk” to CAI classification.
GATE: clinical utility
Strategic boundary. DVJ records one structured injected-fault result; ENABL3S shows that the broader trust story fails: hard loss is detectable, informational faults are mostly missed, and the atlas scarcely exceeds conventional features. More single-corpus simulation tuning is no longer the best path. CAMS answers a separate noncommercial direct-load question; a rights-clean, physically faulted, multi-device Stage 3 determines whether any physical capture-QA capability exists.

13 · Reproduce and inspect

Canonical artifacts

outputs/dvj/measurements.csv
outputs/dvj/trial_accounting.csv
outputs/validation/model_benchmark.csv
outputs/validation/orthogonal_validation.json
outputs/shift_lab/report.json
outputs/capture_qa/report.json
outputs/capture_qa/models.csv
outputs/capture_qa/predictions.csv
outputs/native_two_family/report.json
outputs/native_two_family/trial_metrics.csv
outputs/native_two_family/repeatability.csv
outputs/native_two_family/dose_response.csv
outputs/acquisition/candidate_matrix.csv
outputs/acquisition/decision.json
outputs/acquisition/physionet_admission.json
outputs/physionet/recording_manifest.csv
outputs/physionet/released_clock_audit.csv
outputs/physionet/clock_report.json
outputs/physionet_plurality/report.json
outputs/physionet_plurality/recording_metrics.csv
outputs/physionet_plurality/spectra.csv
outputs/physionet_plurality/speed_matched_permutations.csv
outputs/physionet_local_atlas/report.json
outputs/physionet_atlas_transport/report.json
outputs/physionet_atlas_transport/predictions.csv
outputs/physionet_atlas_transport/same_subject_cross_speed_predictions.csv
outputs/enabl3s_atlas_transport/report.json
outputs/enabl3s_atlas_transport/external_predictions.csv
outputs/enabl3s_atlas_transport/frequency_diagnostics.csv
outputs/enabl3s_atlas_transport/target_native_predictions.csv
outputs/enabl3s_raw_mode_lab/report.json
outputs/enabl3s_raw_mode_lab/stratified_diagnostics.csv
outputs/enabl3s_fault_lab/report.json
outputs/enabl3s_fault_lab/detector_scenarios.csv
outputs/enabl3s_fault_lab/selective_summary.csv
outputs/stage3_capture_protocol/protocol_freeze.json
outputs/instrument_benchmark.json
outputs/research_program_audit.json
outputs/reference_lab_observation.json

Research documents

config/research_program.yaml
config/stage3_capture_protocol.yaml
templates/STAGE3_CAPTURE_MANIFEST_TEMPLATE.csv
templates/STAGE3_PLATFORM_SITE_REGISTER_TEMPLATE.csv
templates/STAGE3_FAULT_SOP_MATRIX_TEMPLATE.csv
templates/STAGE3_REHEARSAL_METRICS_TEMPLATE.csv
docs/COMPUTATIONAL_PHASE_SYNTHESIS.md
docs/PHYSICAL_EXPERIMENT_REQUIREMENTS.md
docs/STAGE3_PHYSICAL_CAPTURE_PROTOCOL.md
docs/REFERENCE_LAB.md
docs/CONSEQUENCE_LAB.md
docs/CAPTURE_QA.md
docs/PHYSIONET_ADMISSION.md
docs/PHYSIONET_PLURALITY.md
docs/LOCAL_PLURALITY_ATLAS.md
docs/ATLAS_TRANSPORT.md
docs/ENABL3S_EXTERNAL_TRANSPORT.md
docs/ENABL3S_RAW_MODE_AND_FAULT_LAB.md
docs/NATIVE_TWO_FAMILY.md
docs/FROZEN_DEVICE_TRANSPORT.md
docs/ACQUISITION_DECISION.md
docs/ORTHOGONAL_INCREMENT.md
docs/INDEPENDENCE_BOUNDARY.md
docs/SHIFT_ABSTENTION.md
docs/COMPETITIVE_POSITION.md
docs/ZTOYBOX_LINEAGE.md
docs/PILOT_RESULTS.md
docs/PREREGISTRATION.md
docs/CLAIMS.md
docs/INDUSTRY_SCORECARD.md
docs/DATA_DICTIONARY.md

Commands
cd /Users/russellparrish/Documents/russell/ZTOYBOX/Biomechanical
PYTHONPATH=src python3 -m pytest tests -q
PYTHONPATH=src python3 -m ztoy_biomech.cli research-audit \
  --program config/research_program.yaml \
  --output outputs/research_program_audit.json
PYTHONPATH=src python3 -m ztoy_biomech.cli instrument-bench \
  --output outputs/instrument_benchmark.json --repetitions 100
PYTHONPATH=src python3 -m ztoy_biomech.cli shift-lab \
  data/extracted/dvj-cai outputs/dvj/measurements.csv \
  --output outputs/shift_lab
PYTHONPATH=src python3 -m ztoy_biomech.cli native-two-family \
  data/extracted/dvj-cai --output outputs/native_two_family
PYTHONPATH=src python3 -m ztoy_biomech.cli capture-qa \
  outputs/shift_lab/shifted_measurements.csv --output outputs/capture_qa
PYTHONPATH=src python3 -m ztoy_biomech.cli acquisition-decision \
  --output outputs/acquisition
PYTHONPATH=src python3 -m ztoy_biomech.cli physionet-audit \
  a-multimodal-gait-dataset-of-brain-activity-muscle-activity-kinematics-and-ground-forces-in-young-adults-1.0.0 \
  --verify-checksums --output outputs/acquisition/physionet_admission.json
PYTHONPATH=src python3 -m ztoy_biomech.cli physionet-clock-audit \
  a-multimodal-gait-dataset-of-brain-activity-muscle-activity-kinematics-and-ground-forces-in-young-adults-1.0.0 \
  --output outputs/physionet
PYTHONPATH=src python3 -m ztoy_biomech.cli physionet-plurality \
  a-multimodal-gait-dataset-of-brain-activity-muscle-activity-kinematics-and-ground-forces-in-young-adults-1.0.0 \
  --output outputs/physionet_plurality --permutations 500
PYTHONPATH=src python3 -m ztoy_biomech.cli physionet-resolution-morphology \
  outputs/physionet_plurality/spectra.csv \
  outputs/physionet_plurality/segment_spectra.csv \
  --output outputs/physionet_resolution_morphology --permutations 500
PYTHONPATH=src python3 -m ztoy_biomech.cli physionet-local-atlas \
  a-multimodal-gait-dataset-of-brain-activity-muscle-activity-kinematics-and-ground-forces-in-young-adults-1.0.0 \
  --output outputs/physionet_local_atlas --permutations 500
PYTHONPATH=src python3 -m ztoy_biomech.cli physionet-atlas-transport \
  outputs/physionet_local_atlas/local_atlas.csv \
  outputs/physionet_resolution_morphology/segment_curves.csv \
  outputs/physionet_local_atlas/segment_features.csv \
  --output outputs/physionet_atlas_transport --bootstrap 1000
PYTHONPATH=src python3 -m ztoy_biomech.cli enabl3s-atlas-transport \
  data/raw/enabl3s outputs/physionet_local_atlas/channel_spectra.csv \
  --output outputs/enabl3s_atlas_transport
PYTHONPATH=src python3 -m ztoy_biomech.cli enabl3s-atlas-transport \
  data/raw/enabl3s outputs/physionet_local_atlas/channel_spectra.csv \
  --sample-rate-hz 500 --output outputs/enabl3s_atlas_transport_clock500
PYTHONPATH=src python3 -m ztoy_biomech.cli enabl3s-raw-mode-lab \
  data/raw/enabl3s --output outputs/enabl3s_raw_mode_lab
PYTHONPATH=src python3 -m ztoy_biomech.cli enabl3s-fault-lab \
  data/raw/enabl3s --output outputs/enabl3s_fault_lab
PYTHONPATH=src python3 -m ztoy_biomech.cli stage3-capture-audit \
  stage3_manifest.csv --data-root /path/to/capture/files \
  --output outputs/stage3_capture_audit
PYTHONPATH=src python3 -m ztoy_biomech.cli stage3-protocol-freeze \
  --protocol config/stage3_capture_protocol.yaml \
  --output outputs/stage3_capture_protocol/protocol_freeze.json
PYTHONPATH=src python3 -m ztoy_biomech.cli reference-lab \
  outputs/dvj/measurements.csv --output outputs/reference-lab
PYTHONPATH=src python3 -m ztoy_biomech.cli consequence-lab \
  waveforms.csv --output outputs/consequence --target internal_load \
  --outcome-kind instrumented_implant --outcome-source implant_sensor \
  --feature-provenance provenance.json --baseline phase force moment \
  --instrument rotation route persistence
PYTHONPATH=src python3 -m ztoy_biomech.cli transport-lab \
  waveforms.csv --output outputs/transport --target internal_load \
  --outcome-kind instrumented_implant --outcome-source implant_sensor \
  --feature-provenance provenance.json --baseline temporal_incumbent \
  --instrument rotation route persistence --device-column device_id
PYTHONPATH=src python3 -m ztoy_biomech.cli dvj-measure \
  data/extracted/dvj-cai --output outputs/dvj
PYTHONPATH=src python3 -m ztoy_biomech.cli dvj-validate \
  outputs/dvj/measurements.csv --output outputs/validation \
  --permutations 100 --bootstrap 1000