Skip to content

Repository files navigation

Mask2Real-SimDataGen

Simulation side of Mask2Real-WM: an Isaac Lab extension (faive_lab) and data-generation pipeline for manipulation tasks with an ORCA hand mounted on a Franka arm. It covers the full loop: record or synthesize demonstrations in simulation, expand them with MimicGen, post-process the resulting HDF5 files, and export everything to LeRobot format. It also renders the simulated ground truth for the Mask2Real-WM controllability evaluation.

Installation

The tested, reproducible install path is the headless installer under scripts/setup/:

bash scripts/setup/install_env.sh     # creates the conda env, installs Isaac Sim/Isaac Lab, ~20-40 min
bash scripts/setup/restore_assets.sh  # recovers gitignored USD/USDZ assets (Git LFS)
source scripts/setup/activate_env.sh  # activate the env, in every new shell
bash scripts/setup/check_env.sh       # verify everything, including gym task registration
  • install_env.sh creates a pinned conda env (Python 3.11, PyTorch 2.7.0+cu128, Isaac Sim 5.0.0, Isaac Lab at a pinned commit) and installs both source packages this repo needs: source/faive_lab and source/faive_data_generation.
  • restore_assets.sh pulls Git LFS content and reports anything still missing.
  • activate_env.sh activates the conda env and sets a couple of environment variables Isaac Sim needs in a headless shell (EULA acceptance, file-descriptor limit).
  • check_env.sh verifies the interpreter, pinned packages, required assets, and — as its final and most important check — actually launches Isaac Sim headless and confirms the orca-synthetic-mimic-ik-abs-v0 gym task registers correctly.

See docs/python_env_setup.md for the reasoning behind every version pin, troubleshooting, and a from-scratch smoke test.

Verify the install manually at any time with:

python scripts/list_envs.py

Repo structure

source/
├── faive_lab/faive_lab/
│   ├── assets/          # robot/object/scene USD assets + their Python configs
│   ├── devices/         # teleoperation device factory + hand-tracking retargeter wrapper
│   ├── retargeter/       # vendored hand-tracking retargeting algorithm (see below)
│   ├── sim/              # custom asset spawners / actuators
│   └── tasks/world_modeling/  # the MimicGen task: orca-synthetic-mimic-ik-abs-v0
└── faive_data_generation/faive_data_generation/datagen/
    # MimicGen-based DataGenerator + the from-scratch RandomGenerator

scripts/
├── setup/            # headless installer (see Installation)
├── data_collection/  # the pipeline scripts documented below
├── wm_evaluation/    # ground-truth renders for the Mask2Real-WM controllability evaluation
├── environments/teleoperation/  # extra teleop device implementations
└── tools/            # HDF5 utilities, mesh conversion, calibration inspection

openxr/   # CloudXR runtime setup for Apple Vision Pro hand-tracking teleop
docs/     # deep-dive setup/troubleshooting docs

Adding new assets

3D assets live under source/faive_lab/faive_lab/assets/data/<category>/. The root .gitignore blanket-ignores all USD types (*.usd, *.usda, *.usdc, *.usdz), and each asset directory has its own .gitignore that re-includes specific filenames — see assets/data/orca_v1/.gitignore for the pattern to copy. Included USD/EXR types are tracked via Git LFS (.gitattributes).

To add a new tracked asset:

  1. Drop the file under assets/data/<category>/.
  2. Add a !<filename> line to that directory's .gitignore (create one if it doesn't exist).
  3. git add the file — it should pick up the LFS filter automatically; confirm with git check-attr filter -- path/to/file (should print filter: lfs).

Environment maps are optional and not part of the repository: every .exr file put into source/faive_lab/faive_lab/assets/data/skyboxes/indoor/ (for example HDRIs from Poly Haven) becomes a candidate texture for the dome light, and one is picked at random each time the task is loaded. Without any, the scene is lit by a plain white dome light.

Quickstart: generating random-motion data

The fastest way to confirm your install works end-to-end, and the simplest way to generate data with no source demonstrations at all, is random-motion generation. It drives the arm/hand with randomized waypoint and joint motion instead of replaying recorded demos:

python scripts/data_collection/collect_synthetic_data.py \
    --task orca-synthetic-mimic-ik-abs-v0 \
    --num_envs 4 \
    --total_num_episodes 8 \
    --output_file datasets/random_motion_smoke_test.hdf5 \
    --headless

The dataset is written next to the given name with a rank suffix, here datasets/random_motion_smoke_test_rank_0.hdf5. The object defaults to banana, which Isaac Sim downloads from NVIDIA's asset server; put ORCA_SPAWN_OBJECT_TYPE=cylinder (or cube, torus) in front of the command to use an object that needs no download.

Cameras must be explicitly enabled or nothing gets recorded. The recorder manager only attaches export terms when RGB, depth, or segmentation cameras are on; otherwise the run reports demos as generated but silently writes an empty dataset. Set ORCA_ENABLE_SEGMENTATION_CAMERAS=true (and pass --enable_cameras) or enable cameras via --config_file, e.g.:

ORCA_ENABLE_SEGMENTATION_CAMERAS=true python scripts/data_collection/collect_synthetic_data.py \
    --task orca-synthetic-mimic-ik-abs-v0 --enable_cameras \
    --num_envs 4 --total_num_episodes 8 \
    --output_file datasets/random_motion_smoke_test.hdf5 --headless

Random-motion behavior is tunable through a --config_file YAML (merged into RandomGeneratorConfig):

config_params:
  drift_strength: 0.6          # how strongly waypoints drift toward the object
  random_noise_scale: 0.05
  sine_hand_motion: true       # sinusoidal motion on a random subset of hand joints
  sine_min_joints: 3
  sine_max_joints: 8
resolution: [320, 240]         # camera resolution override

Recording demos (for MimicGen or standalone)

Recommended: record_and_annotate.py records and annotates each episode with MimicGen subtask signals in a single pass, so its output can be used directly as MimicGen source data — no separate annotation step needed.

python scripts/data_collection/record_and_annotate.py \
    --task orca-synthetic-mimic-ik-abs-v0 \
    --teleop_device keyboard_hand \
    --object_type banana \
    --dataset_file datasets/banana_demos.hdf5 \
    --num_demos 20

Supported --teleop_device values:

Device Notes
keyboard Isaac Lab's built-in Se3 keyboard device.
keyboard_hand Keyboard arm control + preset hand postures (FaiveHandController).
spacemouse 3Dconnexion SpaceMouse.
handtracking Apple Vision Pro via OpenXR/CloudXR, retargeted to the ORCA hand. Requires the one-time CloudXR runtime setup in openxr/Readme.md and XR_RUNTIME_JSON exported before launching.

--object_type accepts dice, cube, sphere, cylinder, cone, torus, duck, banana, cup, mug, or all. duck needs the mesh assets/data/misc/duck_unit.usd, which is not part of the repository.

If you'd rather record and annotate as two separate steps (e.g. to review raw recordings first), use record_demos.py followed by annotate_demos.py — both scripts accept the same --teleop_device/--dataset_file style arguments.

Generating data with MimicGen

Once you have an annotated source dataset (from record_and_annotate.py, or record_demos.py + annotate_demos.py), expand it to many new episodes with MimicGen:

ORCA_ENABLE_SEGMENTATION_CAMERAS=true python scripts/data_collection/collect_synthetic_data.py \
    --task orca-synthetic-mimic-ik-abs-v0 \
    --mimicgen \
    --num_envs 64 \
    --total_num_episodes 1000 \
    --input_file datasets/banana_demos.hdf5 \
    --output_file datasets/generated/banana_mimicgen.hdf5 \
    --dataset_export_mode separate \
    --headless --enable_cameras

--input_file must be annotated (contains MimicGen subtask term/start signals). Feeding it raw record_demos.py output fails with an IndexError deep inside the datagen info pool, because no subtask signal is ever non-zero. The same "cameras must be enabled or nothing is written" rule from the quickstart applies here too.

Omit --mimicgen and the same script falls back to the random-motion generator from the quickstart above — useful for bootstrapping a dataset with no source demos at all.

--config_file also lets you override MimicGen-specific behavior per subtask (mimicgen_subtask_noise, mimicgen_subtask_selection, mimicgen_exclude_demo_ids, etc.) — see scripts/data_collection/collect_synthetic_data.py's argument parsing for the full schema.

Post-processing existing HDF5 files

Two scripts replay an existing HDF5's recorded actions/states through the simulator to add or refresh data, but they differ in cost and what they can change:

rollout_hdf5_demos.py — full physics replay through env.step(), re-exported through the same recorder stack used for collection. Use it when you need to regenerate observations/success labels or change the export mode:

python scripts/data_collection/rollout_hdf5_demos.py \
    --input_file datasets/torus_demos_50.hdf5 \
    --output_file datasets/torus_demos_50_rollout.hdf5 \
    --object_type torus \
    --enable_segmentation_cameras \
    --headless

append_rgb_camera_to_demos.py — a lighter-weight alternative for the common case of adding RGB camera frames to demos that were recorded without cameras (cameras are usually disabled during recording/generation for speed). It kinematically replays only the recorded states snapshots via scene.reset_to(...) — no physics stepping, no drift risk — purely to render and inject rgb_camera/{wrist_camera,side_camera_one} groups into the HDF5, in place or into a copy:

python scripts/data_collection/append_rgb_camera_to_demos.py \
    --input_file datasets/torus_demos_50.hdf5 \
    --object_type torus \
    --headless

Use rollout_hdf5_demos.py when you need a genuine re-simulation; use append_rgb_camera_to_demos.py when you just need RGB frames added quickly.

Converting to LeRobot format

Two scripts, not interchangeable — they read different HDF5 schemas:

  • convert_to_lerobot.py — for HDF5 produced by this repo's simulation pipeline (actions/processed_actions + obs/ + rgb_camera/segmentation_camera/ depth_camera groups from the Isaac Lab recorder stack).
  • real_convert_to_lerobot.py — for HDF5 captured on real hardware (actions_arm/actions_hand + observations/qpos_arm/qpos_hand + observations/images/<camera>/color, with fixed OAK-D camera names).
# Simulation data
python scripts/data_collection/convert_to_lerobot.py \
    --input datasets/banana_demos_rollout.hdf5 \
    --output datasets/lerobot/banana \
    --num-workers 4 --fps 25

# Real-robot data
python scripts/data_collection/real_convert_to_lerobot.py \
    --input datasets/real_banana.hdf5 \
    --output datasets/lerobot_real/banana \
    --fps 30

Both write the same output layout (annotation/<episode_id>.json + videos/<episode_id>/<view>_rgb.mp4), both run in the lightweight environment-convert-to-lerobot.yml conda env (no Isaac Sim required — just h5py, scipy, and mediapy), and both send an optional Discord notification on completion via DISCORD_WEBHOOK_URL in a .env file (silently skipped if unset).

Ground-truth renders for the Mask2Real-WM controllability evaluation

The controllability evaluation of Mask2Real-WM asks a world model to reach target poses of the arm and hand and compares its last frame with what reaching the target really looks like. scripts/wm_evaluation/render_controllability_targets.py produces that reference: for every target it holds the start pose and then the target pose as a setpoint until the simulated robot has settled, and renders both cameras.

python scripts/wm_evaluation/render_controllability_targets.py \
    --targets_manifest /path/to/Mask2Real-WM/outputs/controllability/combined_targets_manifest.json \
    --output_dir outputs/gt_renders \
    --headless --enable_cameras

--targets_manifest is the file written by the targets step of the Mask2Real-WM evaluation (see its docs/controllability_eval.md). The script writes, per target,

outputs/gt_renders/<trial_group_id>/
    <camera>_rgb.png               RGB render
    <camera>_seg_preview.png       false-color preview of the instance segmentation
    <camera>_seg_raw_ids.npy       instance id of every pixel
    <camera>_id_to_labels.json     prim path of every instance id
outputs/gt_renders/gt_manifest.json

for the cameras side_camera_one and wrist_camera. gt_manifest.json lists, per target, the render paths, the instance ids that belong to the hand and to the arm, and the settle diagnostics (converged, the remaining joint error, and saturated_dims for joints that did not reach the target). Running the script again with the same --output_dir skips targets that are already rendered; --max_targets N renders a subset.

Notes:

  • The default --object_type cylinder uses a local asset. banana, cup and mug are fetched from a remote server and are not needed here.
  • The Franka arm is hidden in the data-generation task and shown again by this script, so that it can be segmented separately from the hand.
  • scripts/wm_evaluation/verify_bridge_against_real_sample.py drives the simulation to the last state of a real recording and saves the render, as a check that commanded poses are reached and that the cameras match the real ones.

Camera calibration

The side and wrist cameras are placed and parametrized from a calibration of the real camera rig, which is loaded in orca_synthetic_data_gen_mimic_env_cfg.py:

CALIBRATION_PARAMS_PATH = {
    "transformations": f"{FAIVE_ASSETS_DATA_DIR}/calibration_params_08_03_26/transformations.pkl",
    "camera_intrinsics": f"{FAIVE_ASSETS_DATA_DIR}/calibration_params_08_03_26/camera_intrinsics.pkl",
}
CAMERA_NAMES = ["oakd_side_view", "oakd_wrist_view"]

The repository ships the calibration of the rig used for Mask2Real-WM in source/faive_lab/faive_lab/assets/data/calibration_params_08_03_26/. The task needs these two files; the simulated cameras then match the real side and wrist cameras.

Pickle formats:

  • camera_intrinsics.pkl: dict[camera_name -> (intrinsic_matrix[3,3], distortion_coeffs[14])], keyed by the names in CAMERA_NAMES.
  • transformations.pkl: list[(camera_name, extrinsic_4x4)]. Extrinsics are looked up by list position, not by name — the Nth entry must correspond to CAMERA_NAMES[N]. This is a common source of silently-swapped cameras; double-check ordering if a camera ends up looking from the wrong pose.

To use another rig, produce matching pickles from your own calibration, put them under source/faive_lab/faive_lab/assets/data/<your_calibration_dir>/, and update CALIBRATION_PARAMS_PATH/CAMERA_NAMES above. Use scripts/tools/read_calibration_params.ipynb to inspect a pickle before wiring it in.

Code formatting

pip install pre-commit
pre-commit run --all-files

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages