Simulation side of Mask2Real-WM: an Isaac Lab
extension (faive_lab) and data-generation pipeline for manipulation tasks with an ORCA
hand mounted on a Franka arm. It covers the full loop: record or synthesize
demonstrations in simulation, expand them with MimicGen,
post-process the resulting HDF5 files, and export everything to
LeRobot format. It also renders the simulated
ground truth for the Mask2Real-WM controllability evaluation.
The tested, reproducible install path is the headless installer under
scripts/setup/:
bash scripts/setup/install_env.sh # creates the conda env, installs Isaac Sim/Isaac Lab, ~20-40 min
bash scripts/setup/restore_assets.sh # recovers gitignored USD/USDZ assets (Git LFS)
source scripts/setup/activate_env.sh # activate the env, in every new shell
bash scripts/setup/check_env.sh # verify everything, including gym task registrationinstall_env.shcreates a pinned conda env (Python 3.11, PyTorch 2.7.0+cu128, Isaac Sim 5.0.0, Isaac Lab at a pinned commit) and installs both source packages this repo needs:source/faive_labandsource/faive_data_generation.restore_assets.shpulls Git LFS content and reports anything still missing.activate_env.shactivates the conda env and sets a couple of environment variables Isaac Sim needs in a headless shell (EULA acceptance, file-descriptor limit).check_env.shverifies the interpreter, pinned packages, required assets, and — as its final and most important check — actually launches Isaac Sim headless and confirms theorca-synthetic-mimic-ik-abs-v0gym task registers correctly.
See docs/python_env_setup.md for the reasoning behind
every version pin, troubleshooting, and a from-scratch smoke test.
Verify the install manually at any time with:
python scripts/list_envs.pysource/
├── faive_lab/faive_lab/
│ ├── assets/ # robot/object/scene USD assets + their Python configs
│ ├── devices/ # teleoperation device factory + hand-tracking retargeter wrapper
│ ├── retargeter/ # vendored hand-tracking retargeting algorithm (see below)
│ ├── sim/ # custom asset spawners / actuators
│ └── tasks/world_modeling/ # the MimicGen task: orca-synthetic-mimic-ik-abs-v0
└── faive_data_generation/faive_data_generation/datagen/
# MimicGen-based DataGenerator + the from-scratch RandomGenerator
scripts/
├── setup/ # headless installer (see Installation)
├── data_collection/ # the pipeline scripts documented below
├── wm_evaluation/ # ground-truth renders for the Mask2Real-WM controllability evaluation
├── environments/teleoperation/ # extra teleop device implementations
└── tools/ # HDF5 utilities, mesh conversion, calibration inspection
openxr/ # CloudXR runtime setup for Apple Vision Pro hand-tracking teleop
docs/ # deep-dive setup/troubleshooting docs
3D assets live under source/faive_lab/faive_lab/assets/data/<category>/. The root
.gitignore blanket-ignores all USD types (*.usd, *.usda, *.usdc, *.usdz), and
each asset directory has its own .gitignore that re-includes specific filenames —
see assets/data/orca_v1/.gitignore
for the pattern to copy. Included USD/EXR types are tracked via Git LFS
(.gitattributes).
To add a new tracked asset:
- Drop the file under
assets/data/<category>/. - Add a
!<filename>line to that directory's.gitignore(create one if it doesn't exist). git addthe file — it should pick up the LFS filter automatically; confirm withgit check-attr filter -- path/to/file(should printfilter: lfs).
Environment maps are optional and not part of the repository: every .exr file put into
source/faive_lab/faive_lab/assets/data/skyboxes/indoor/ (for example HDRIs from
Poly Haven) becomes a candidate texture for the dome light,
and one is picked at random each time the task is loaded. Without any, the scene is lit by
a plain white dome light.
The fastest way to confirm your install works end-to-end, and the simplest way to generate data with no source demonstrations at all, is random-motion generation. It drives the arm/hand with randomized waypoint and joint motion instead of replaying recorded demos:
python scripts/data_collection/collect_synthetic_data.py \
--task orca-synthetic-mimic-ik-abs-v0 \
--num_envs 4 \
--total_num_episodes 8 \
--output_file datasets/random_motion_smoke_test.hdf5 \
--headlessThe dataset is written next to the given name with a rank suffix, here
datasets/random_motion_smoke_test_rank_0.hdf5. The object defaults to banana, which
Isaac Sim downloads from NVIDIA's asset server; put ORCA_SPAWN_OBJECT_TYPE=cylinder (or
cube, torus) in front of the command to use an object that needs no download.
Cameras must be explicitly enabled or nothing gets recorded. The recorder manager only attaches export terms when RGB, depth, or segmentation cameras are on; otherwise the run reports demos as generated but silently writes an empty dataset. Set
ORCA_ENABLE_SEGMENTATION_CAMERAS=true(and pass--enable_cameras) or enable cameras via--config_file, e.g.:ORCA_ENABLE_SEGMENTATION_CAMERAS=true python scripts/data_collection/collect_synthetic_data.py \ --task orca-synthetic-mimic-ik-abs-v0 --enable_cameras \ --num_envs 4 --total_num_episodes 8 \ --output_file datasets/random_motion_smoke_test.hdf5 --headless
Random-motion behavior is tunable through a --config_file YAML (merged into
RandomGeneratorConfig):
config_params:
drift_strength: 0.6 # how strongly waypoints drift toward the object
random_noise_scale: 0.05
sine_hand_motion: true # sinusoidal motion on a random subset of hand joints
sine_min_joints: 3
sine_max_joints: 8
resolution: [320, 240] # camera resolution overrideRecommended: record_and_annotate.py records and annotates each episode with
MimicGen subtask signals in a single pass, so its output can be used directly as
MimicGen source data — no separate annotation step needed.
python scripts/data_collection/record_and_annotate.py \
--task orca-synthetic-mimic-ik-abs-v0 \
--teleop_device keyboard_hand \
--object_type banana \
--dataset_file datasets/banana_demos.hdf5 \
--num_demos 20Supported --teleop_device values:
| Device | Notes |
|---|---|
keyboard |
Isaac Lab's built-in Se3 keyboard device. |
keyboard_hand |
Keyboard arm control + preset hand postures (FaiveHandController). |
spacemouse |
3Dconnexion SpaceMouse. |
handtracking |
Apple Vision Pro via OpenXR/CloudXR, retargeted to the ORCA hand. Requires the one-time CloudXR runtime setup in openxr/Readme.md and XR_RUNTIME_JSON exported before launching. |
--object_type accepts dice, cube, sphere, cylinder, cone, torus, duck,
banana, cup, mug, or all. duck needs the mesh assets/data/misc/duck_unit.usd,
which is not part of the repository.
If you'd rather record and annotate as two separate steps (e.g. to review raw
recordings first), use record_demos.py followed by annotate_demos.py — both
scripts accept the same --teleop_device/--dataset_file style arguments.
Once you have an annotated source dataset (from record_and_annotate.py, or
record_demos.py + annotate_demos.py), expand it to many new episodes with
MimicGen:
ORCA_ENABLE_SEGMENTATION_CAMERAS=true python scripts/data_collection/collect_synthetic_data.py \
--task orca-synthetic-mimic-ik-abs-v0 \
--mimicgen \
--num_envs 64 \
--total_num_episodes 1000 \
--input_file datasets/banana_demos.hdf5 \
--output_file datasets/generated/banana_mimicgen.hdf5 \
--dataset_export_mode separate \
--headless --enable_cameras
--input_filemust be annotated (contains MimicGen subtask term/start signals). Feeding it rawrecord_demos.pyoutput fails with anIndexErrordeep inside the datagen info pool, because no subtask signal is ever non-zero. The same "cameras must be enabled or nothing is written" rule from the quickstart applies here too.
Omit --mimicgen and the same script falls back to the random-motion generator from
the quickstart above — useful for bootstrapping a dataset with no source demos at all.
--config_file also lets you override MimicGen-specific behavior per subtask
(mimicgen_subtask_noise, mimicgen_subtask_selection, mimicgen_exclude_demo_ids,
etc.) — see scripts/data_collection/collect_synthetic_data.py's argument parsing for
the full schema.
Two scripts replay an existing HDF5's recorded actions/states through the simulator to add or refresh data, but they differ in cost and what they can change:
rollout_hdf5_demos.py — full physics replay through env.step(), re-exported
through the same recorder stack used for collection. Use it when you need to
regenerate observations/success labels or change the export mode:
python scripts/data_collection/rollout_hdf5_demos.py \
--input_file datasets/torus_demos_50.hdf5 \
--output_file datasets/torus_demos_50_rollout.hdf5 \
--object_type torus \
--enable_segmentation_cameras \
--headlessappend_rgb_camera_to_demos.py — a lighter-weight alternative for the common case
of adding RGB camera frames to demos that were recorded without cameras (cameras are
usually disabled during recording/generation for speed). It kinematically replays only
the recorded states snapshots via scene.reset_to(...) — no physics stepping, no
drift risk — purely to render and inject rgb_camera/{wrist_camera,side_camera_one}
groups into the HDF5, in place or into a copy:
python scripts/data_collection/append_rgb_camera_to_demos.py \
--input_file datasets/torus_demos_50.hdf5 \
--object_type torus \
--headlessUse rollout_hdf5_demos.py when you need a genuine re-simulation; use
append_rgb_camera_to_demos.py when you just need RGB frames added quickly.
Two scripts, not interchangeable — they read different HDF5 schemas:
convert_to_lerobot.py— for HDF5 produced by this repo's simulation pipeline (actions/processed_actions+obs/+rgb_camera/segmentation_camera/depth_cameragroups from the Isaac Lab recorder stack).real_convert_to_lerobot.py— for HDF5 captured on real hardware (actions_arm/actions_hand+observations/qpos_arm/qpos_hand+observations/images/<camera>/color, with fixed OAK-D camera names).
# Simulation data
python scripts/data_collection/convert_to_lerobot.py \
--input datasets/banana_demos_rollout.hdf5 \
--output datasets/lerobot/banana \
--num-workers 4 --fps 25
# Real-robot data
python scripts/data_collection/real_convert_to_lerobot.py \
--input datasets/real_banana.hdf5 \
--output datasets/lerobot_real/banana \
--fps 30Both write the same output layout (annotation/<episode_id>.json +
videos/<episode_id>/<view>_rgb.mp4), both run in the lightweight
environment-convert-to-lerobot.yml conda env (no Isaac Sim required — just h5py,
scipy, and mediapy), and both send an optional Discord notification on completion via
DISCORD_WEBHOOK_URL in a .env file (silently skipped if unset).
The controllability evaluation of
Mask2Real-WM asks a world model to reach
target poses of the arm and hand and compares its last frame with what reaching the
target really looks like. scripts/wm_evaluation/render_controllability_targets.py
produces that reference: for every target it holds the start pose and then the target
pose as a setpoint until the simulated robot has settled, and renders both cameras.
python scripts/wm_evaluation/render_controllability_targets.py \
--targets_manifest /path/to/Mask2Real-WM/outputs/controllability/combined_targets_manifest.json \
--output_dir outputs/gt_renders \
--headless --enable_cameras--targets_manifest is the file written by the targets step of the Mask2Real-WM
evaluation (see its docs/controllability_eval.md). The script writes, per target,
outputs/gt_renders/<trial_group_id>/
<camera>_rgb.png RGB render
<camera>_seg_preview.png false-color preview of the instance segmentation
<camera>_seg_raw_ids.npy instance id of every pixel
<camera>_id_to_labels.json prim path of every instance id
outputs/gt_renders/gt_manifest.json
for the cameras side_camera_one and wrist_camera. gt_manifest.json lists, per
target, the render paths, the instance ids that belong to the hand and to the arm, and
the settle diagnostics (converged, the remaining joint error, and saturated_dims
for joints that did not reach the target). Running the script again with the same
--output_dir skips targets that are already rendered; --max_targets N renders a
subset.
Notes:
- The default
--object_type cylinderuses a local asset.banana,cupandmugare fetched from a remote server and are not needed here. - The Franka arm is hidden in the data-generation task and shown again by this script, so that it can be segmented separately from the hand.
scripts/wm_evaluation/verify_bridge_against_real_sample.pydrives the simulation to the last state of a real recording and saves the render, as a check that commanded poses are reached and that the cameras match the real ones.
The side and wrist cameras are placed and parametrized from a calibration of the real
camera rig, which is loaded in
orca_synthetic_data_gen_mimic_env_cfg.py:
CALIBRATION_PARAMS_PATH = {
"transformations": f"{FAIVE_ASSETS_DATA_DIR}/calibration_params_08_03_26/transformations.pkl",
"camera_intrinsics": f"{FAIVE_ASSETS_DATA_DIR}/calibration_params_08_03_26/camera_intrinsics.pkl",
}
CAMERA_NAMES = ["oakd_side_view", "oakd_wrist_view"]The repository ships the calibration of the rig used for Mask2Real-WM in
source/faive_lab/faive_lab/assets/data/calibration_params_08_03_26/. The task needs
these two files; the simulated cameras then match the real side and wrist cameras.
Pickle formats:
camera_intrinsics.pkl:dict[camera_name -> (intrinsic_matrix[3,3], distortion_coeffs[14])], keyed by the names inCAMERA_NAMES.transformations.pkl:list[(camera_name, extrinsic_4x4)]. Extrinsics are looked up by list position, not by name — the Nth entry must correspond toCAMERA_NAMES[N]. This is a common source of silently-swapped cameras; double-check ordering if a camera ends up looking from the wrong pose.
To use another rig, produce matching pickles from your own calibration, put them under
source/faive_lab/faive_lab/assets/data/<your_calibration_dir>/, and update
CALIBRATION_PARAMS_PATH/CAMERA_NAMES above. Use
scripts/tools/read_calibration_params.ipynb
to inspect a pickle before wiring it in.
pip install pre-commit
pre-commit run --all-files