This repository contains the minimal long-video Phase 1 tuning workspace for the dual-sampling pipeline.
phase1_longvideo/: standalone long-video sampling packagesplit/: train/val/test split manifestsscripts/: metadata build, subset creation, evaluation, sweep, and validation utilitiescolab_phase1_longvideo.ipynb: Colab workflow for the long-video tuning loop
The long-video dataset is not committed to git. Place the extracted dataset shards in a local dataset/ directory before running metadata build.
Phase 1 baseline is uniform sampling with the same final_num_frames budget as the dual-stage run.
Run separate baselines for 16, 32, and 64 frames.
- Build metadata:
python scripts/build_phase1_longvideo_metadata.py- Create a small subset:
python scripts/sample_phase1_subset.py `
--subset-name smoke12 `
--target-video-count 12 `
--source-split train- Run a quick synthetic smoke evaluation:
python scripts/eval_phase1_sampling.py `
--input results/subsets/smoke12/smoke12_train_metadata.jsonl `
--split train `
--backend synthetic `
--limit 20- Run a real SigLIP sweep in Colab:
python scripts/sweep_phase1_sampling.py `
--input results/subsets/sample40/sample40_train_metadata.jsonl `
--subset-name sample40 `
--split train `
--backend siglip `
--siglip-device cuda `
--resume- Prefer
T4first for this Phase 1 sweep. - Switch to
L4only if aT4run is too slow. - Avoid
A100for this round; it is better reserved for LoRA or end-to-end runs. - Copy the active video subset from Drive to local Colab disk before large sweeps to reduce Drive random I/O overhead.
For group-level comparisons with Phase 2, standardize on:
mIoUR@1@0.3R@1@0.5R@1@0.7BERTScoreAvgRuntimeSec
Do not treat BLEU/ROUGE/METEOR from inference_test.py as the main shared table metrics.