V-Cartridges is a video question-answering benchmark evaluation framework. We take a set of videos and benchmark questions, synthesize compressed "cartridge" representations of each video using a language model, and then train lightweight adapters on those cartridges before running evaluation. This pipeline lets you run large-scale video QA experiments efficiently, with support for MLVU and VideoMME benchmarks, both locally and on Modal. Generated artifacts are logged to runs/ and results/.
conda create -n vcartridges python=3.11
conda activate vcartridges
pip install uv
uv pip install -e ".[dev,modal]"
vcart setup-checkPlace datasets under:
data/mlvu/
data/videomme/
Run the full experiment:
python scripts/reproduce_vcart_results.py \
--spec experiments/vcart_results_100.toml \
--run-dir runs/vcart-results-100 \
--phase all \
--resumeRun individual stages:
vcart freeze-items --spec experiments/vcart_results_100.toml --run-dir runs/vcart-results-100
vcart synthesize-cartridges --spec experiments/vcart_results_100.toml --run-dir runs/vcart-results-100 --resume
vcart train-cartridges --spec experiments/vcart_results_100.toml --run-dir runs/vcart-results-100 --resume
vcart eval-matrix --spec experiments/vcart_results_100.toml --run-dir runs/vcart-results-100 --resume
vcart aggregate-results --spec experiments/vcart_results_100.toml --run-dir runs/vcart-results-100Preview a run without executing it:
python scripts/reproduce_vcart_results.py \
--spec experiments/vcart_results_100.toml \
--run-dir runs/vcart-results-100 \
--phase all \
--dry-runRestrict a run with:
--only-model <model>
--only-benchmark <benchmark>
--only-video-id <video_id>
--limit <count>
Create the Modal resources:
modal volume create vcartridges-runs
modal secret create vcartridges-secrets WANDB_API_KEY=<key> HF_TOKEN=<token>Upload MLVU data:
modal volume put vcartridges-runs /local/path/to/MLVU /data/mlvuRun the experiment:
modal run --timestamps modal_mlvu.py::prepare_data --benchmark mlvu
modal run --timestamps modal_mlvu.py::prepare_data --benchmark videomme
modal run --timestamps modal_mlvu.py::freeze_items
modal run --timestamps modal_mlvu.py::synthesize_cartridges
modal run --timestamps modal_mlvu.py::train_cartridges
modal run --timestamps modal_mlvu.py::eval_matrix
modal run --timestamps modal_mlvu.py::aggregate_resultsDownload the generated artifacts:
modal volume get vcartridges-runs /runs/vcart-results-100 ./runs/vcart-results-100