Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VLA Manipulation: Ablations, Generalization, and a Constrained-VRAM Study

Empirical study of vision-language-action (VLA) policies for a simulated cube-stacking task in NVIDIA Isaac Lab, run entirely on a 6 GB laptop GPU (RTX 4050 Laptop, ~5.64 GiB usable). Includes a from-scratch behavior-cloning policy, a 7-arm ablation grid, a held-out generalization campaign, and a hardware feasibility study of loading OpenVLA-7B on the same machine.

Every number in this repo comes from a script in scripts/ and a result file in data/. Nothing here is a prediction. Several were, and were wrong, before being replaced by measurement.

What's actually here

A behavior-cloning VLA policy (scripts/vla_model.py, scripts/train_vla.py). Vision + language + proprioceptive state → gripper action, trained on ~100 teleop-style demonstrations per target color.

A 7-arm ablation grid (scripts/ablate_vla.py): baseline, unfrozen vision encoder, frame-split augmentation, ResNet-50 backbone, no-language input, no-FiLM conditioning, state-only (vision-blind). Each arm evaluated over 24 in-distribution rollouts.

A held-out generalization campaign (scripts/register_heldout_task.py, scripts/eval_heldout_frames.py): the same 7 checkpoints re-evaluated on cube spawn positions that do not overlap the training distribution at all, built as an Isaac Lab task config subclass rather than assumed from the registration call.

An OpenVLA-7B feasibility study (scripts/openvla_infer.py, scripts/probe_vram.py, scripts/read_safetensors_header.py): can a 7.5B-parameter VLA foundation model run inference on a 6 GB card, and if not, exactly where does it fail.

Statistical tooling (scripts/fisher_exact.py, scripts/test_distance_sweep_significance.py, scripts/compare_heldout_vs_indist.py): Fisher's exact test for low-count binary outcomes, Mann-Whitney U for continuous per-episode metrics, applied consistently rather than picking whichever test makes a result look better.

Findings

1. The vision encoder is not the memory constraint. Every encoder candidate tested (ResNet-18/34/50, ViT-S/B, EfficientNet-B0) fits in under 0.8 GB forward-pass memory (data/vram_probe.json). Freezing the encoder in the main policy is a decision forced by data (~100 demonstrations cannot usefully train a 22M-parameter vision stack), not by hardware.

2. OpenVLA-7B fails four times before VRAM is ever the problem. In order: a _supports_sdpa property accessed before assignment under transformers>=4.50; a hard timm version assertion in the model's own remote code; a huggingface-hub version conflict introduced by pinning timm; a system pandas/numpy ABI mismatch. Only after all four are resolved does the model reach an actual load attempt, which then fails predictably in fp16 (OOM during weight loading, not the forward pass) and succeeds in int4. Parameter decomposition, read directly from the safetensors header over HTTP range requests (no 15 GB download required): the language model is 89.4% of the 7.54B parameters; the vision-to-language projector, arguably the actual "VLA" part, is under 1% (data/openvla_params.json).

stack rate by evaluation regime

Every bar is measured. Flat ticks mark a measured zero, so a rate of 0.0 is visibly distinct from a condition that was never evaluated. Regenerate with python3 tools/make_figures.py data/camerashift_summary.json docs/generalization_gap.gif.

3. In-distribution, every ablation arm scored zero cube-stacks across 168 rollouts. The main ablation campaign (7 arms × 24 episodes) produced 0.0 stack rate for every single arm on in-distribution cube positions.

4. On held-out cube positions, stacking got better, not worse. The same 7 checkpoints, evaluated on spawn positions with zero overlap with the training distribution, produced 8 total stacks across the same 168-rollout budget. The harder condition outperformed the easier one (data/heldout_vs_indist_comparison.json). Of 7 per-arm significance tests (Fisher's exact, held-out vs. in-distribution grasp counts), exactly one clears p<0.05: the no_language arm went from the best in-distribution grasper (7/24) to zero held-out grasps (p=0.009), consistent with, but not proof of, a vision-only shortcut that breaks once cube positions move outside the training range. Reported as a single significant result out of seven tests, with the ~30% chance of at least one false positive at that sample size stated alongside it, not hidden.

5. A null result at one statistical resolution was a real effect at another. A distance-sweep experiment found 0 of 9 grasp/stack-count comparisons significant via Fisher's exact test. Re-testing the same rollouts' raw gripper-switch-rate traces (continuous, not pooled to an integer) with Mann-Whitney U found 4 of 9 significant, all three within the unfrozen arm, with the baseline arm's 0-of-3 result acting as a control (data/distance_sweep_gripper_comparison.json). Same data, same sample size. resolving the metric at the trace level rather than the outcome-count level recovered signal that pooling had thrown away. This does not retract the original null result for the metric it tested; it separates "no effect" from "the instrument couldn't see it."

6. Behavior-cloned gripper actions are lagged relative to their own labels, but the lag is not why the task fails. A lag-sweep against validation frames finds peak sign-accuracy at lag +1, not lag 0 (99.0% vs. 97.4%), and near-perfect sign accuracy (up to 100%) when aligned to gripper flip events specifically (data/gripper_lag.json). That reads like the obvious cause of the zero stack rate in finding 3. It isn't: replaying the live closed-loop scripted expert with its gripper channel artificially delayed by 0-3 steps costs nothing at all: 16/16 stacks at every delay up to 3, and still 16/16 at 8 (data/arm_tolerance.json, data/delay_causal_wide.json). Only a 20-step delay (roughly 20x the measured policy lag) collapses the task. The one-step delay is real and precisely characterized; it is also an order of magnitude too small to explain the failure.

7. What actually explains the failure: the correlation of the policy's error, not its magnitude. Five sequential hypotheses about the gripper and the arm were each tested and each falsified in turn (badly-fit gripper channel, hedged gripper output, one-step delay, i.i.d. arm-position noise at the policy's own measured error scale, all cost nothing). The experiment that finally moved the needle held the magnitude of injected arm-position noise fixed at the policy's own per-step error (σ=0.02) and varied only its temporal correlation, using a live closed-loop control arm (100% stack, confirming the harness itself is not the confound) and a sensitivity arm large enough to guarantee failure (σ=0.08 → 0% stack, confirming the harness can detect an effect):

perturbation (σ=0.02, same magnitude throughout) lag-1 autocorrelation stack rate grasp rate
i.i.d. (independent draw each step) 0.00 87.5% 100.0%
drift (Ornstein-Uhlenbeck, τ=40) 0.975 62.5% 100.0%
bias (one offset held for the whole episode) 1.00 37.5% 31.2%

Monotone in autocorrelation at constant magnitude, in three roughly equal 25-point steps (data/arm_correlation.json). Independent per-step noise mostly cancels over a trajectory; a held or slowly-drifting offset of the same size does not, and compounds into a position error the policy never corrects, which is exactly what a trained policy's own error looks like (the same visual state produces the same mistake every time it recurs), unlike the i.i.d. noise typically used to test robustness. The magnitude of a policy's error was never the interesting variable; whether that error repeats is.

Hardware

RTX 4050 Laptop GPU, 6.05 GB total / 5.64 GiB usable under PyTorch. Every result above was produced on this card; the OpenVLA feasibility study exists specifically because a 7.5B-parameter model does not fit that budget without int4 quantization.

Repo scope

This repo contains the modeling, ablation, evaluation, and analysis code plus the measured result files (JSON) and small per-episode trace arrays (NumPy). Trained checkpoints and raw demonstration recordings (multi-GB video/state captures) are not included, the result files are the record of what those runs produced.

Figures

tools/make_figures.py regenerates the README figure from the result files in data/, so a figure cannot drift from the numbers behind it:

python3 -m pip install numpy matplotlib pillow
python3 tools/make_figures.py data/camerashift_summary.json docs/generalization_gap.gif

License

MIT

About

VLA manipulation ablations, held-out generalization, and an OpenVLA-7B feasibility study, all measured on a 6GB laptop GPU

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages