You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
pipeline.init.intrinsics=gt is unreachable from vipe infer: the video and frame-directory paths cannot ingest a user calibration #105
The typed configuration exposes pipeline.init.intrinsics=gt, and slam.optimize_intrinsics already
resolves to false when it is selected. But neither vipe infer video.mp4 nor vipe infer --image-dir
has a way to hand a calibration to the pipeline: RawMp4Stream and FrameDirStream never populate frame
intrinsics, and DefaultAnnotationPipeline unconditionally appends GeoCalib and asserts that no
intrinsics are present. Selecting gt from the CLI therefore fails for every user who already knows
their camera, which is the common case for rendered, rig-calibrated or benchmark footage (#37, #56).
A companion PR adds --intrinsics calibration.json to both input paths, with regression tests. The
rest of this issue is the controlled evidence for why the option matters, gathered with that patch
applied, under a standard Brown radial distortion with every evaluated pixel real content.
On the same frames, the stock pinhole-plus-GeoCalib path still reports 100.0% registration under the distortion while its ATE RMSE is 1.3x to 17.2x that of the undistorted frames, depending on the scene; undistorting recovers most of that, and supplying the fixed true intrinsics changes the undistorted ATE by a factor between 0.8x and 3.0x (B over C), so it helps on some scenes and not on others. None of the three predeclared mechanism thresholds was met (see the results section for why). Everything below is rendered from hash-sealed run receipts.
Calibrated input: cell C runs through the companion patch (vipe-1.2.0-calibrated-frame-dir.patch),
which adds --intrinsics calibration.json and lets DefaultAnnotationPipeline honour init.intrinsics=gt. Cells A and B use the stock GeoCalib path and are unaffected by the patch.
The companion PR carries only the calibration wiring, rebased onto current main with the video
path wired as well; the --seed flag and the resolved-configuration record that the runs here
used stay in this patch and are not proposed upstream.
Every run writes the Hydra-resolved configuration next to its outputs; a cell is accepted only when the
resolved pipeline.init.intrinsics / pipeline.slam.optimize_intrinsics pair matches its condition.
Camera model and normalization
Rendered pinhole camera; the distortion applied for condition A is opencv-brown-pinhole in camera-normalized coordinates with k1=-0.15, k2=0.0, p1=0.0, p2=0.0, k3=0.0; border black,
resampling bilinear. Condition B inverts it exactly on the inner monotonic branch. Every condition is centre-cropped to the largest window in which every A and B pixel is real content (policy center-crop-to-common-valid-region, 2 px margin); the preparer refuses a frame with any invalid pixel inside the window.
Relation to Experimental: stabilize MEI intrinsics optimization #103: the earlier k1 = -0.3 material behind that PR used a distortion operator normalised
by the image half-diagonal, not camera-normalised Brown/OpenCV, as its author found. This suite uses the
standard model, which is why the coefficient and the results differ from that PR's tables.
Ground-truth pinhole intrinsics supplied to C (fx = fy, principal point at the image centre):
Dataset
Size
fx
fy
cx
cy
gt-human-trajectory-v2
1652x930
1158.0337
1158.0337
826.0
465.0
gt-vfc-texture-rich-v1
1710x960
1303.6753
1303.6753
855.0
480.0
gt-vfc-texture-poor-v1
1710x960
1303.6753
1303.6753
855.0
480.0
Conditions
Cell
Pixels
Intrinsics
Optimize intrinsics
A
brown-distorted
geocalib
yes
B
correctly-undistorted-from-A
geocalib
yes
C
byte-identical-to-B
ground-truth-pinhole
no
Each cell runs on a Live clip (moving performers) and a Frozen clip (same camera path, performers frozen at
one source instant): 180 frames each, 3 rendered environments, matched seed 20260826 for every invocation.
Random seeds
--seed 20260826 seeds Python, NumPy and torch RNGs in every run.
The pinned ViPE CUDA implementation contains a documented nondeterministic scatter path; a seed controls declared RNG state but does not make the pipeline deterministic. This 18-run suite is a single-seed screening experiment. Repeat it under separately predeclared seeds before making a general stability claim.
Results
Suite vipe-brown-multi-environment-v2, 18 runs, matched seed 20260826, determinism claim seed-controlled-not-deterministic.
ATE RMSE and orientation error are computed per clip after a Sim(3) (Umeyama) alignment of the estimated camera centres onto the exact authored truth, over that clip's registered frames, so scale is not scored. Registered is the share of the 180 frames for which ViPE emitted a pose. The same definitions apply to every suite below.
gt-human-trajectory-v2 — South Boston fight gym (not offered publicly)
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.0446
1.229
100.0
A
frozen
0.0392
1.286
100.0
B
live
0.0026
0.291
100.0
B
frozen
0.0025
0.219
100.0
C
live
0.0031
0.177
100.0
C
frozen
0.0026
0.276
100.0
gt-vfc-texture-rich-v1 — Vertical Fight City, texture-rich path (public kit)
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.0923
2.247
100.0
A
frozen
0.0559
2.534
100.0
B
live
0.0344
2.273
100.0
B
frozen
0.0168
1.280
100.0
C
live
0.0114
1.821
100.0
C
frozen
0.0112
0.980
100.0
gt-vfc-texture-poor-v1 — Vertical Fight City, texture-poor path
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.0329
2.948
100.0
A
frozen
0.0188
1.644
100.0
B
live
0.0253
2.254
100.0
B
frozen
0.0150
1.136
100.0
C
live
0.0156
1.500
100.0
C
frozen
0.0144
1.084
100.0
Mechanism findings (predeclared thresholds, applied by the scorer)
Hypothesis
Status
Support
Threshold
Camera-model mismatch (A→B)
NOT_SUPPORTED
0 clip comparisons
≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
GeoCalib / intrinsics optimization (B→C)
NOT_SUPPORTED
0 clip comparisons
≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
Dynamic foreground / masking (Live vs Frozen in C)
NOT_SUPPORTED
0 environments
≥2 environments with Live ATE ≥25% and ≥0.03 m above Frozen
Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.
The orientation criterion (a reduction of at least 3.0°) was predeclared for a regime in which the distorted condition misorients the camera by many degrees. Where a cell is already below that floor, the criterion cannot be met whatever the ATE change, so NOT_SUPPORTED there reads as "threshold not reached", not as "no effect". The thresholds were left as predeclared rather than tuned after the fact.
Suite vipe-brown-multi-environment-v1, 18 runs, matched seed 20260826, determinism claim seed-controlled-not-deterministic.
gt-human-trajectory-v2 — South Boston fight gym (not offered publicly)
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.1453
8.807
100.0
A
frozen
0.1203
6.874
100.0
B
live
0.0040
0.378
100.0
B
frozen
0.0038
0.438
100.0
C
live
0.0017
0.129
100.0
C
frozen
0.0014
0.235
100.0
gt-vfc-texture-rich-v1 — Vertical Fight City, texture-rich path (public kit)
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.2574
21.072
100.0
A
frozen
0.1944
12.157
100.0
B
live
0.0128
1.145
100.0
B
frozen
0.0085
0.632
100.0
C
live
0.0121
1.244
100.0
C
frozen
0.0067
0.561
100.0
gt-vfc-texture-poor-v1 — Vertical Fight City, texture-poor path
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.2427
45.100
100.0
A
frozen
0.0200
2.569
100.0
B
live
0.0154
1.636
100.0
B
frozen
0.0110
0.808
100.0
C
live
0.0127
1.727
100.0
C
frozen
0.0112
0.935
100.0
Mechanism findings (predeclared thresholds, applied by the scorer)
Hypothesis
Status
Support
Threshold
Camera-model mismatch (A→B)
SUPPORTED
5 clip comparisons
≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
GeoCalib / intrinsics optimization (B→C)
NOT_SUPPORTED
0 clip comparisons
≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
Dynamic foreground / masking (Live vs Frozen in C)
NOT_SUPPORTED
0 environments
≥2 environments with Live ATE ≥25% and ≥0.03 m above Frozen
Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.
Superseded by the primary suite above. Measured 2026-09-10 on the sealed v1 inputs: 26.5% (Vertical Fight City) and 32.5% (South Boston) of every A frame is black, the corners and side edges entirely, while B carries none. v1's camera-model-mismatch finding therefore mixes the wrong lens model with a third of the image missing. v2 keeps every A/B/C pixel as real content. The contracts differ in: distortion, randomness, suiteId. Reported for completeness; read the primary.
Replication of the superseded variant under a second seed: vipe-brown-multi-environment-v1-seed-replication-1
Suite vipe-brown-multi-environment-v1-seed-replication-1, 18 runs, matched seed 20260904, determinism claim seed-controlled-not-deterministic.
gt-human-trajectory-v2 — South Boston fight gym (not offered publicly)
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.0998
2.887
100.0
A
frozen
0.1159
6.724
100.0
B
live
0.0040
0.368
100.0
B
frozen
0.0038
0.435
100.0
C
live
0.0017
0.129
100.0
C
frozen
0.0014
0.235
100.0
gt-vfc-texture-rich-v1 — Vertical Fight City, texture-rich path (public kit)
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.2509
20.574
100.0
A
frozen
0.1910
11.811
100.0
B
live
0.0136
1.234
100.0
B
frozen
0.0085
0.638
100.0
C
live
0.0120
1.244
100.0
C
frozen
0.0068
0.559
100.0
gt-vfc-texture-poor-v1 — Vertical Fight City, texture-poor path
Cell
Clip
ATE RMSE (m)
Orientation mean (deg)
Registered (%)
A
live
0.2768
45.066
100.0
A
frozen
0.0202
2.610
100.0
B
live
0.0154
1.640
100.0
B
frozen
0.0109
0.809
100.0
C
live
0.0126
1.748
100.0
C
frozen
0.0112
0.934
100.0
Mechanism findings (predeclared thresholds, applied by the scorer)
Hypothesis
Status
Support
Threshold
Camera-model mismatch (A→B)
SUPPORTED
4 clip comparisons
≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
GeoCalib / intrinsics optimization (B→C)
NOT_SUPPORTED
0 clip comparisons
≥4 comparisons with ≥25% ATE reduction and ≥3.0° orientation reduction
Dynamic foreground / masking (Live vs Frozen in C)
NOT_SUPPORTED
0 environments
≥2 environments with Live ATE ≥25% and ≥0.03 m above Frozen
Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.
Replicates the superseded variant above, not the primary: same inputs as that variant, seed 20260904, predeclared 2026-09-04. Reported for completeness.
Diagnostic reading (predeclared)
A fails and B,C work consistently across environments → supports camera-model-mismatch
B fails while C works consistently across environments → supports geocalib-or-intrinsics-optimization
Live remains worse than Frozen in C within an environment → supports dynamic-foreground-or-masking
None of these outcomes establishes a general ViPE defect, a physical-lens claim, or held-out generalization.
Environment
gpuName: NVIDIA A100-SXM4-80GB
computeCapability: 8.0
torch: 2.13.0+cu130
environment lock: pip freeze --all of the execution interpreter, sha256 b75e862eb5bc89adcc83de3548e4ac45a4015baf94a6c20c98def1bb399061f7
Exact inputs (the evaluated PNGs for A and B/C, hash-sealed), the pinhole calibration, the answer key, the
patch, and the runner/scorer used here, for the Vertical Fight City texture-rich environment. Built by build_vipe_brown_reproduction_kit.py; hosted on Zenodo as gt-vfc-texture-rich-v1.zip (777 entries, 1,114,635,730 bytes, SHA-256 9e7c19fe5776be37993b3623b2e753f3deeb3f9b932ab9d39eadd5789a01f16a), DOI 10.5281/zenodo.22753017. Published on Zenodo 2026-09-14; the DOI resolves to the public record. The archive is large because it ships the
evaluated PNGs themselves, hash-sealed, rather than a renderer to regenerate them; nothing in it needs
the source capture.
The generally applicable issue
The typed configuration exposes
pipeline.init.intrinsics=gt, andslam.optimize_intrinsicsalreadyresolves to
falsewhen it is selected. But neithervipe infer video.mp4norvipe infer --image-dirhas a way to hand a calibration to the pipeline:
RawMp4StreamandFrameDirStreamnever populate frameintrinsics, and
DefaultAnnotationPipelineunconditionally appends GeoCalib and asserts that nointrinsics are present. Selecting
gtfrom the CLI therefore fails for every user who already knowstheir camera, which is the common case for rendered, rig-calibrated or benchmark footage (#37, #56).
A companion PR adds
--intrinsics calibration.jsonto both input paths, with regression tests. Therest of this issue is the controlled evidence for why the option matters, gathered with that patch
applied, under a standard Brown radial distortion with every evaluated pixel real content.
On the same frames, the stock pinhole-plus-GeoCalib path still reports 100.0% registration under the distortion while its ATE RMSE is 1.3x to 17.2x that of the undistorted frames, depending on the scene; undistorting recovers most of that, and supplying the fixed true intrinsics changes the undistorted ATE by a factor between 0.8x and 3.0x (B over C), so it helps on some scenes and not on others. None of the three predeclared mechanism thresholds was met (see the results section for why). Everything below is rendered from hash-sealed run receipts.
ViPE version and configuration
1.2.0, source commit95a8816947602ddc26fcb7a80bea4f9313059578(https://github.com/nv-tlabs/vipe)default(configs/pipeline/default.yaml),init.camera_type=pinholevipe-1.2.0-calibrated-frame-dir.patch),which adds
--intrinsics calibration.jsonand letsDefaultAnnotationPipelinehonourinit.intrinsics=gt. Cells A and B use the stock GeoCalib path and are unaffected by the patch.The companion PR carries only the calibration wiring, rebased onto current
mainwith the videopath wired as well; the
--seedflag and the resolved-configuration record that the runs hereused stay in this patch and are not proposed upstream.
resolved
pipeline.init.intrinsics/pipeline.slam.optimize_intrinsicspair matches its condition.Camera model and normalization
opencv-brown-pinholeincamera-normalizedcoordinates withk1=-0.15, k2=0.0, p1=0.0, p2=0.0, k3=0.0; borderblack,resampling
bilinear. Condition B inverts it exactly on the inner monotonic branch. Every condition is centre-cropped to the largest window in which every A and B pixel is real content (policycenter-crop-to-common-valid-region, 2 px margin); the preparer refuses a frame with any invalid pixel inside the window.by the image half-diagonal, not camera-normalised Brown/OpenCV, as its author found. This suite uses the
standard model, which is why the coefficient and the results differ from that PR's tables.
gt-human-trajectory-v2gt-vfc-texture-rich-v1gt-vfc-texture-poor-v1Conditions
Each cell runs on a Live clip (moving performers) and a Frozen clip (same camera path, performers frozen at
one source instant): 180 frames each, 3 rendered environments, matched seed
20260826for every invocation.Random seeds
--seed 20260826seeds Python, NumPy and torch RNGs in every run.The pinned ViPE CUDA implementation contains a documented nondeterministic scatter path; a seed controls declared RNG state but does not make the pipeline deterministic. This 18-run suite is a single-seed screening experiment. Repeat it under separately predeclared seeds before making a general stability claim.
Results
Suite
vipe-brown-multi-environment-v2, 18 runs, matched seed20260826, determinism claimseed-controlled-not-deterministic.ATE RMSE and orientation error are computed per clip after a Sim(3) (Umeyama) alignment of the estimated camera centres onto the exact authored truth, over that clip's registered frames, so scale is not scored. Registered is the share of the 180 frames for which ViPE emitted a pose. The same definitions apply to every suite below.
gt-human-trajectory-v2— South Boston fight gym (not offered publicly)gt-vfc-texture-rich-v1— Vertical Fight City, texture-rich path (public kit)gt-vfc-texture-poor-v1— Vertical Fight City, texture-poor pathMechanism findings (predeclared thresholds, applied by the scorer)
Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.
The orientation criterion (a reduction of at least 3.0°) was predeclared for a regime in which the distorted condition misorients the camera by many degrees. Where a cell is already below that floor, the criterion cannot be met whatever the ATE change, so NOT_SUPPORTED there reads as "threshold not reached", not as "no effect". The thresholds were left as predeclared rather than tuned after the fact.
Earlier variant, superseded:
vipe-brown-multi-environment-v1Suite
vipe-brown-multi-environment-v1, 18 runs, matched seed20260826, determinism claimseed-controlled-not-deterministic.gt-human-trajectory-v2— South Boston fight gym (not offered publicly)gt-vfc-texture-rich-v1— Vertical Fight City, texture-rich path (public kit)gt-vfc-texture-poor-v1— Vertical Fight City, texture-poor pathMechanism findings (predeclared thresholds, applied by the scorer)
Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.
Superseded by the primary suite above. Measured 2026-09-10 on the sealed v1 inputs: 26.5% (Vertical Fight City) and 32.5% (South Boston) of every A frame is black, the corners and side edges entirely, while B carries none. v1's camera-model-mismatch finding therefore mixes the wrong lens model with a third of the image missing. v2 keeps every A/B/C pixel as real content. The contracts differ in:
distortion,randomness,suiteId. Reported for completeness; read the primary.Replication of the superseded variant under a second seed:
vipe-brown-multi-environment-v1-seed-replication-1Suite
vipe-brown-multi-environment-v1-seed-replication-1, 18 runs, matched seed20260904, determinism claimseed-controlled-not-deterministic.gt-human-trajectory-v2— South Boston fight gym (not offered publicly)gt-vfc-texture-rich-v1— Vertical Fight City, texture-rich path (public kit)gt-vfc-texture-poor-v1— Vertical Fight City, texture-poor pathMechanism findings (predeclared thresholds, applied by the scorer)
Single-seed thresholded findings diagnose these three fixed rendered environments; they do not establish run-to-run stability or a general ViPE defect.
Replicates the superseded variant above, not the primary: same inputs as that variant, seed
20260904, predeclared 2026-09-04. Reported for completeness.Diagnostic reading (predeclared)
None of these outcomes establishes a general ViPE defect, a physical-lens claim, or held-out generalization.
Environment
NVIDIA A100-SXM4-80GB8.02.13.0+cu130pip freeze --allof the execution interpreter, sha256b75e862eb5bc89adcc83de3548e4ac45a4015baf94a6c20c98def1bb399061f7numpy==2.5.2opencv-python==5.0.0.93torch==2.13.0+cu130torchvision==0.28.0+cu130gcpa2-ultragpu-1g/evercoast-eval-a100/us-central1-a1080Reproduction package
Exact inputs (the evaluated PNGs for A and B/C, hash-sealed), the pinhole calibration, the answer key, the
patch, and the runner/scorer used here, for the Vertical Fight City texture-rich environment. Built by
build_vipe_brown_reproduction_kit.py; hosted on Zenodo asgt-vfc-texture-rich-v1.zip(777 entries, 1,114,635,730 bytes, SHA-2569e7c19fe5776be37993b3623b2e753f3deeb3f9b932ab9d39eadd5789a01f16a), DOI 10.5281/zenodo.22753017. Published on Zenodo 2026-09-14; the DOI resolves to the public record. The archive is large because it ships theevaluated PNGs themselves, hash-sealed, rather than a renderer to regenerate them; nothing in it needs
the source capture.
Sources this text was rendered from
vipe-brown-calibrated-ablation-v2.jsonsha256aabe4581408e7629e99e7f3995b30f6b4cadfd849cc45f87b66ea4b85a2ad7dcd6a31e514db2febb5e9b81c9183e6b7a7015d3369585d6cd7258085d7fa6d437db2dbcf910826cb8f51eb711dd3728e60d34fe472f52117de373e147f0626570vipe-brown-calibrated-ablation-v1.jsonsha2566b437ab0ba98e8e90ed81b4823b8ce0aa83fe95ada58dd3be98ef15fa59d09b1, comparison sha25610c29b31a0fe0bcacded99569c299aaf5298158eb1cb62421f940ee8960a4c14vipe-brown-calibrated-ablation-v1-seed-replication-1.jsonsha256eac156618356fb390b47f1060d9f71ecdef2640ee39703400bb037bea96f2f2f, comparison sha256700fecd871c9438c472a7bf1e1217e76d29c619bb8b7a03b65cf8b2834223ad6gt-vfc-texture-rich-v1: 776 files, 1115.5 MB, manifest sha2561373b1caf1aa59e327cdae5fd9c3d13a381c9a5843a74ed79bc4e7bddaaf6314KIT-HOSTING.jsonsha256568363a207d68d009ecca47f2ad2a7e02ccdd9c9640795e07fa243253676e499