Skip to content

Repository files navigation

shift-safe-monitors

Failure, Not Novelty: Disentangled and Adaptively Calibrated Runtime Monitors for Robot Policies Under Distribution Shift.

Full research proposal: PROPOSAL.md.

Layout

src/ssm/
  traces.py            # episode trace schema + parquet IO (collection<->analysis contract)
  signals.py           # the five monitor signal families, unified per-step interface
  conformal.py         # calibration ladder: split / weighted / delayed-feedback ACI / few-shot
  metrics.py           # AUROC, ECE, coverage, time-to-detection, confusion decomposition
  dismon.py            # two-head disentangled monitor (rungs A/B/C), torch
  stats.py             # bootstrap CIs, paired permutation tests, Holm-Bonferroni
  plotting.py          # all five paper figures from tidy DataFrames
  release.py           # HF dataset packaging: shards + integrity manifest + card
  experiments/
    pipeline.py        # Part 1: benchmark under shift (+ cross-policy transfer)
    ladder.py          # Part 3: L0-L3 recalibration engine
    dismon_protocol.py # Part 2: held-out-dimension invariance protocol + p-values
  collect/
    adapters.py        # policy adapters (pi0.5, OpenVLA-OFT interfaces)
    hooks.py           # FeatureTap + ActivationPerturber (tested torch instrumentation)
    runner.py          # checkpoint-resumable sweep runner
    libero_env.py      # LIBERO-Plus env wrapper (chunk stepping locally tested)
    mock.py            # simulator-free mock world for full-scale local dry runs
colab/
  quickstart.ipynb     # day-1 preflight + smoke test (COMPLETE 2026-07-10, all gates green)
  stages/              # stage0 dryrun, stage1 collect, stage2 smolvla finetune,
                       # stage3 analyze, stage4 dismon+ladder, stage5 figures
tests/                 # 52 local tests incl. end-to-end pipeline + collection integration
REPRODUCE.md           # figure-by-figure reproduction commands + statistical reporting rules

Try the whole pipeline right now (no GPU, no simulator)

uv run python colab/stages/stage0_dryrun.py --out dryrun
# ~75 s: collects 1,120 mock episodes through the real manifests/runner, runs the full
# analysis (benchmark, DisMon, calibration ladder), renders all 5 figures, packages a
# release. Output under dryrun/ is clearly MOCK data with the paper's expected structure.

Local dev

uv sync            # core + dev deps (torch CPU for dismon tests)
uv run pytest -q   # all tests must pass before any sweep
uv run ruff check .

Colab

Open colab/quickstart.ipynb, follow the cells in order. Stage scripts in colab/stages/ are checkpoint-resumable: every sweep writes per-episode parquet traces and can be interrupted/resumed at any time.

Data contract

Every rollout, clean or perturbed, becomes one EpisodeTrace (see src/ssm/traces.py): per-step monitor scores and pooled policy features, plus episode metadata (policy, suite, task, perturbation dimension, severity, outcome). All analysis code consumes only traces, so collection and analysis can proceed independently.

About

Safety monitors for learned robot policies under distribution shift. A benchmark for which monitor still works once the deployment distribution moves.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages