Failure, Not Novelty: Disentangled and Adaptively Calibrated Runtime Monitors for Robot Policies Under Distribution Shift.
Full research proposal: PROPOSAL.md.
src/ssm/
traces.py # episode trace schema + parquet IO (collection<->analysis contract)
signals.py # the five monitor signal families, unified per-step interface
conformal.py # calibration ladder: split / weighted / delayed-feedback ACI / few-shot
metrics.py # AUROC, ECE, coverage, time-to-detection, confusion decomposition
dismon.py # two-head disentangled monitor (rungs A/B/C), torch
stats.py # bootstrap CIs, paired permutation tests, Holm-Bonferroni
plotting.py # all five paper figures from tidy DataFrames
release.py # HF dataset packaging: shards + integrity manifest + card
experiments/
pipeline.py # Part 1: benchmark under shift (+ cross-policy transfer)
ladder.py # Part 3: L0-L3 recalibration engine
dismon_protocol.py # Part 2: held-out-dimension invariance protocol + p-values
collect/
adapters.py # policy adapters (pi0.5, OpenVLA-OFT interfaces)
hooks.py # FeatureTap + ActivationPerturber (tested torch instrumentation)
runner.py # checkpoint-resumable sweep runner
libero_env.py # LIBERO-Plus env wrapper (chunk stepping locally tested)
mock.py # simulator-free mock world for full-scale local dry runs
colab/
quickstart.ipynb # day-1 preflight + smoke test (COMPLETE 2026-07-10, all gates green)
stages/ # stage0 dryrun, stage1 collect, stage2 smolvla finetune,
# stage3 analyze, stage4 dismon+ladder, stage5 figures
tests/ # 52 local tests incl. end-to-end pipeline + collection integration
REPRODUCE.md # figure-by-figure reproduction commands + statistical reporting rules
uv run python colab/stages/stage0_dryrun.py --out dryrun
# ~75 s: collects 1,120 mock episodes through the real manifests/runner, runs the full
# analysis (benchmark, DisMon, calibration ladder), renders all 5 figures, packages a
# release. Output under dryrun/ is clearly MOCK data with the paper's expected structure.uv sync # core + dev deps (torch CPU for dismon tests)
uv run pytest -q # all tests must pass before any sweep
uv run ruff check .Open colab/quickstart.ipynb, follow the cells in order. Stage scripts in colab/stages/
are checkpoint-resumable: every sweep writes per-episode parquet traces and can be
interrupted/resumed at any time.
Every rollout, clean or perturbed, becomes one EpisodeTrace (see src/ssm/traces.py):
per-step monitor scores and pooled policy features, plus episode metadata
(policy, suite, task, perturbation dimension, severity, outcome). All analysis code
consumes only traces, so collection and analysis can proceed independently.