Skip to content

pipeline.init.intrinsics=gt is unreachable from vipe infer: the video and frame-directory paths cannot ingest a user calibration #105

Description

@pb-evercoast

The generally applicable issue

The typed configuration exposes pipeline.init.intrinsics=gt, and slam.optimize_intrinsics already
resolves to false when it is selected. But neither vipe infer video.mp4 nor vipe infer --image-dir
has a way to hand a calibration to the pipeline: RawMp4Stream and FrameDirStream never populate frame
intrinsics, and DefaultAnnotationPipeline unconditionally appends GeoCalib and asserts that no
intrinsics are present. Selecting gt from the CLI therefore fails for every user who already knows
their camera, which is the common case for rendered, rig-calibrated or benchmark footage (#37, #56).

A companion PR adds --intrinsics calibration.json to both input paths, with regression tests. The
rest of this issue is the controlled evidence for why the option matters, gathered with that patch
applied, under a standard Brown radial distortion with every evaluated pixel real content.
On the same frames, the stock pinhole-plus-GeoCalib path still reports 100.0% registration under the distortion while its ATE RMSE is 1.3x to 17.2x that of the undistorted frames, depending on the scene; undistorting recovers most of that, and supplying the fixed true intrinsics changes the undistorted ATE by a factor between 0.8x and 3.0x (B over C), so it helps on some scenes and not on others. None of the three predeclared mechanism thresholds was met (see the results section for why). Everything below is rendered from hash-sealed run receipts.

ViPE version and configuration

  • ViPE 1.2.0, source commit 95a8816947602ddc26fcb7a80bea4f9313059578 (https://github.com/nv-tlabs/vipe)
  • Pipeline: default (configs/pipeline/default.yaml), init.camera_type=pinhole
  • Calibrated input: cell C runs through the companion patch (vipe-1.2.0-calibrated-frame-dir.patch),
    which adds --intrinsics calibration.json and lets DefaultAnnotationPipeline honour
    init.intrinsics=gt. Cells A and B use the stock GeoCalib path and are unaffected by the patch.
    The companion PR carries only the calibration wiring, rebased onto current main with the video
    path wired as well; the --seed flag and the resolved-configuration record that the runs here
    used stay in this patch and are not proposed upstream.
  • Every run writes the Hydra-resolved configuration next to its outputs; a cell is accepted only when the
    resolved pipeline.init.intrinsics / pipeline.slam.optimize_intrinsics pair matches its condition.

Camera model and normalization

  • Rendered pinhole camera; the distortion applied for condition A is opencv-brown-pinhole in
    camera-normalized coordinates with k1=-0.15, k2=0.0, p1=0.0, p2=0.0, k3=0.0; border black,
    resampling bilinear. Condition B inverts it exactly on the inner monotonic branch. Every condition is centre-cropped to the largest window in which every A and B pixel is real content (policy center-crop-to-common-valid-region, 2 px margin); the preparer refuses a frame with any invalid pixel inside the window.
  • Relation to Experimental: stabilize MEI intrinsics optimization #103: the earlier k1 = -0.3 material behind that PR used a distortion operator normalised
    by the image half-diagonal, not camera-normalised Brown/OpenCV, as its author found. This suite uses the
    standard model, which is why the coefficient and the results differ from that PR's tables.
  • Ground-truth pinhole intrinsics supplied to C (fx = fy, principal point at the image centre):
Dataset Size fx fy cx cy
gt-human-trajectory-v2 1652x930 1158.0337 1158.0337 826.0 465.0
gt-vfc-texture-rich-v1 1710x960 1303.6753 1303.6753 855.0 480.0
gt-vfc-texture-poor-v1 1710x960 1303.6753 1303.6753 855.0 480.0

Conditions

Cell Pixels Intrinsics Optimize intrinsics
A brown-distorted geocalib yes
B correctly-undistorted-from-A geocalib yes
C byte-identical-to-B ground-truth-pinhole no

Each cell runs on a Live clip (moving performers) and a Frozen clip (same camera path, performers frozen at
one source instant): 180 frames each, 3 rendered environments, matched seed
20260826 for every invocation.

Random seeds

--seed 20260826 seeds Python, NumPy and torch RNGs in every run.
The pinned ViPE CUDA implementation contains a documented nondeterministic scatter path; a seed controls declared RNG state but does not make the pipeline deterministic. This 18-run suite is a single-seed screening experiment. Repeat it under separately predeclared seeds before making a general stability claim.

Results

Suite vipe-brown-multi-environment-v2, 18 runs, matched seed 20260826, determinism claim seed-controlled-not-deterministic.

ATE RMSE and orientation error are computed per clip after a Sim(3) (Umeyama) alignment of the estimated camera centres onto the exact authored truth, over that clip's registered frames, so scale is not scored. Registered is the share of the 180 frames for which ViPE emitted a pose. The same definitions apply to every suite below.

gt-human-trajectory-v2 — South Boston fight gym (not offered publicly)

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.0446 1.229 100.0
A frozen 0.0392 1.286 100.0
B live 0.0026 0.291 100.0
B frozen 0.0025 0.219 100.0
C live 0.0031 0.177 100.0
C frozen 0.0026 0.276 100.0

gt-vfc-texture-rich-v1 — Vertical Fight City, texture-rich path (public kit)

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.0923 2.247 100.0
A frozen 0.0559 2.534 100.0
B live 0.0344 2.273 100.0
B frozen 0.0168 1.280 100.0
C live 0.0114 1.821 100.0
C frozen 0.0112 0.980 100.0

gt-vfc-texture-poor-v1 — Vertical Fight City, texture-poor path

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.0329 2.948 100.0
A frozen 0.0188 1.644 100.0
B live 0.0253 2.254 100.0
B frozen 0.0150 1.136 100.0
C live 0.0156 1.500 100.0
C frozen 0.0144 1.084 100.0

Mechanism findings (predeclared thresholds, applied by the scorer)

Hypothesis Status Support Threshold
Camera-model mismatch (A→B) NOT_SUPPORTED 0 clip comparisons ≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
GeoCalib / intrinsics optimization (B→C) NOT_SUPPORTED 0 clip comparisons ≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
Dynamic foreground / masking (Live vs Frozen in C) NOT_SUPPORTED 0 environments ≥2 environments with Live ATE ≥25% and ≥0.03 m above Frozen

Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.

The orientation criterion (a reduction of at least 3.0°) was predeclared for a regime in which the distorted condition misorients the camera by many degrees. Where a cell is already below that floor, the criterion cannot be met whatever the ATE change, so NOT_SUPPORTED there reads as "threshold not reached", not as "no effect". The thresholds were left as predeclared rather than tuned after the fact.

Earlier variant, superseded: vipe-brown-multi-environment-v1

Suite vipe-brown-multi-environment-v1, 18 runs, matched seed 20260826, determinism claim seed-controlled-not-deterministic.

gt-human-trajectory-v2 — South Boston fight gym (not offered publicly)

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.1453 8.807 100.0
A frozen 0.1203 6.874 100.0
B live 0.0040 0.378 100.0
B frozen 0.0038 0.438 100.0
C live 0.0017 0.129 100.0
C frozen 0.0014 0.235 100.0

gt-vfc-texture-rich-v1 — Vertical Fight City, texture-rich path (public kit)

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.2574 21.072 100.0
A frozen 0.1944 12.157 100.0
B live 0.0128 1.145 100.0
B frozen 0.0085 0.632 100.0
C live 0.0121 1.244 100.0
C frozen 0.0067 0.561 100.0

gt-vfc-texture-poor-v1 — Vertical Fight City, texture-poor path

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.2427 45.100 100.0
A frozen 0.0200 2.569 100.0
B live 0.0154 1.636 100.0
B frozen 0.0110 0.808 100.0
C live 0.0127 1.727 100.0
C frozen 0.0112 0.935 100.0

Mechanism findings (predeclared thresholds, applied by the scorer)

Hypothesis Status Support Threshold
Camera-model mismatch (A→B) SUPPORTED 5 clip comparisons ≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
GeoCalib / intrinsics optimization (B→C) NOT_SUPPORTED 0 clip comparisons ≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
Dynamic foreground / masking (Live vs Frozen in C) NOT_SUPPORTED 0 environments ≥2 environments with Live ATE ≥25% and ≥0.03 m above Frozen

Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.

Superseded by the primary suite above. Measured 2026-09-10 on the sealed v1 inputs: 26.5% (Vertical Fight City) and 32.5% (South Boston) of every A frame is black, the corners and side edges entirely, while B carries none. v1's camera-model-mismatch finding therefore mixes the wrong lens model with a third of the image missing. v2 keeps every A/B/C pixel as real content. The contracts differ in: distortion, randomness, suiteId. Reported for completeness; read the primary.

Replication of the superseded variant under a second seed: vipe-brown-multi-environment-v1-seed-replication-1

Suite vipe-brown-multi-environment-v1-seed-replication-1, 18 runs, matched seed 20260904, determinism claim seed-controlled-not-deterministic.

gt-human-trajectory-v2 — South Boston fight gym (not offered publicly)

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.0998 2.887 100.0
A frozen 0.1159 6.724 100.0
B live 0.0040 0.368 100.0
B frozen 0.0038 0.435 100.0
C live 0.0017 0.129 100.0
C frozen 0.0014 0.235 100.0

gt-vfc-texture-rich-v1 — Vertical Fight City, texture-rich path (public kit)

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.2509 20.574 100.0
A frozen 0.1910 11.811 100.0
B live 0.0136 1.234 100.0
B frozen 0.0085 0.638 100.0
C live 0.0120 1.244 100.0
C frozen 0.0068 0.559 100.0

gt-vfc-texture-poor-v1 — Vertical Fight City, texture-poor path

Cell Clip ATE RMSE (m) Orientation mean (deg) Registered (%)
A live 0.2768 45.066 100.0
A frozen 0.0202 2.610 100.0
B live 0.0154 1.640 100.0
B frozen 0.0109 0.809 100.0
C live 0.0126 1.748 100.0
C frozen 0.0112 0.934 100.0

Mechanism findings (predeclared thresholds, applied by the scorer)

Hypothesis Status Support Threshold
Camera-model mismatch (A→B) SUPPORTED 4 clip comparisons ≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
GeoCalib / intrinsics optimization (B→C) NOT_SUPPORTED 0 clip comparisons ≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
Dynamic foreground / masking (Live vs Frozen in C) NOT_SUPPORTED 0 environments ≥2 environments with Live ATE ≥25% and ≥0.03 m above Frozen

Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.

Replicates the superseded variant above, not the primary: same inputs as that variant, seed 20260904, predeclared 2026-09-04. Reported for completeness.

Diagnostic reading (predeclared)

  • A fails and B,C work consistently across environments → supports camera-model-mismatch
  • B fails while C works consistently across environments → supports geocalib-or-intrinsics-optimization
  • Live remains worse than Frozen in C within an environment → supports dynamic-foreground-or-masking

None of these outcomes establishes a general ViPE defect, a physical-lens claim, or held-out generalization.

Environment

  • gpuName: NVIDIA A100-SXM4-80GB
  • computeCapability: 8.0
  • torch: 2.13.0+cu130
  • environment lock: pip freeze --all of the execution interpreter, sha256 b75e862eb5bc89adcc83de3548e4ac45a4015baf94a6c20c98def1bb399061f7
    • numpy==2.5.2
    • opencv-python==5.0.0.93
    • torch==2.13.0+cu130
    • torchvision==0.28.0+cu130
  • provider: gcp
  • machine: a2-ultragpu-1g/evercoast-eval-a100/us-central1-a
  • wallTimeCapMinutes: 1080

Reproduction package

Exact inputs (the evaluated PNGs for A and B/C, hash-sealed), the pinhole calibration, the answer key, the
patch, and the runner/scorer used here, for the Vertical Fight City texture-rich environment. Built by
build_vipe_brown_reproduction_kit.py; hosted on Zenodo as gt-vfc-texture-rich-v1.zip (777 entries, 1,114,635,730 bytes, SHA-256 9e7c19fe5776be37993b3623b2e753f3deeb3f9b932ab9d39eadd5789a01f16a), DOI 10.5281/zenodo.22753017. Published on Zenodo 2026-09-14; the DOI resolves to the public record. The archive is large because it ships the
evaluated PNGs themselves, hash-sealed, rather than a renderer to regenerate them; nothing in it needs
the source capture.

Sources this text was rendered from

  • contract vipe-brown-calibrated-ablation-v2.json sha256 aabe4581408e7629e99e7f3995b30f6b4cadfd849cc45f87b66ea4b85a2ad7dc
  • calibrated-input patch sha256 d6a31e514db2febb5e9b81c9183e6b7a7015d3369585d6cd7258085d7fa6d437
  • suite comparison sha256 db2dbcf910826cb8f51eb711dd3728e60d34fe472f52117de373e147f0626570
  • supplementary contract vipe-brown-calibrated-ablation-v1.json sha256 6b437ab0ba98e8e90ed81b4823b8ce0aa83fe95ada58dd3be98ef15fa59d09b1, comparison sha256 10c29b31a0fe0bcacded99569c299aaf5298158eb1cb62421f940ee8960a4c14
  • supplementary contract vipe-brown-calibrated-ablation-v1-seed-replication-1.json sha256 eac156618356fb390b47f1060d9f71ecdef2640ee39703400bb037bea96f2f2f, comparison sha256 700fecd871c9438c472a7bf1e1217e76d29c619bb8b7a03b65cf8b2834223ad6
  • public kit gt-vfc-texture-rich-v1: 776 files, 1115.5 MB, manifest sha256 1373b1caf1aa59e327cdae5fd9c3d13a381c9a5843a74ed79bc4e7bddaaf6314
  • kit hosting record KIT-HOSTING.json sha256 568363a207d68d009ecca47f2ad2a7e02ccdd9c9640795e07fa243253676e499

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions