diff --git a/README.md b/README.md index 7325df4..4d0caae 100644 --- a/README.md +++ b/README.md @@ -7,275 +7,215 @@ ██║ ╚██████╔╝██║ ██║╚██████╔╝███████╗ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝ -

⚒ Robotics Data Toolkit ⚒

-Convert, inspect, visualize, score, and discover robotics datasets across every major format. +

⚒ Robotics Data Toolkit & Data Engine ⚒

+Convert between every major robotics format — then turn a pile of datasets into a queryable, searchable, curatable corpus.

PyPI Open in Colab Website Python 3.10+ License: MIT -

-RLDS ═══╗ ╔═══► LeRobot
-HDF5 ═══╣ ╠═══► MCAP
-Zarr ═══╬════⚙════╬═══► RoboDM
-MCAP ═══╝ ╚═══► RLDS

-Convert between robotics dataset formats with one command. Score demonstration quality with research-backed metrics. Lint datasets for hygiene defects before training. Segment episodes into sub-skills with changepoint detection. +Forge is two things that share one core: -| Format | Read | Write | Visualize | Notes | -|--------|:----:|:-----:|:---------:|-------| -| RLDS | ✓ | ✓ | ✓ | Open-X, TensorFlow Datasets | -| LeRobot v2/v3 | ✓ | ✓ | ✓ | HuggingFace, Parquet + MP4 | -| GR00T | ✓ | - | ✓ | NVIDIA Isaac, LeRobot v2 with embodiment metadata | -| RoboDM | ✓ | ✓ | ✓ | Berkeley's .vla format, up to 70x compression* | -| Zarr | ✓ | - | ✓ | Diffusion Policy, UMI | -| HDF5 | ✓ | - | ✓ | robomimic, ACT/ALOHA | -| MCAP | ✓ | ✓ | ✓ | ROS2 CDR + Foxglove Protobuf, no ROS install required | -| Rosbag | ✓ | - | ✓ | ROS1 .bag, ROS2 SQLite3 | +- **A format toolkit** — convert, inspect, score, lint, filter, segment, and visualize a single dataset across RLDS, LeRobot, HDF5, MCAP, Zarr, Rosbag, and more. +- **A data engine** — register every episode you collect into an append-only **catalog**, then query it with SQL, search it by natural language, dedup and curate it, and explore it in a visual **Studio**. -*\*RoboDM requires manual installation from GitHub (see below)* +Everything works on local paths, `hf://` datasets, and `s3://` / `gs://` buckets. -See [docs/model_formats.md](docs/model_formats.md) for which models (Octo, OpenVLA, ACT, Diffusion Policy, etc.) use which format. See [docs/format_reference.md](docs/format_reference.md) for detailed format specifications. + + + + +
-## Why Forge? +**Toolkit**  ·  [Install](#install) · [Convert](#convert--interop) · [Quality](#score--clean) · [Filter](#score--clean) · [Segment](#understand) · [Tokenize](#understand) · [Visualize](#understand) -Every robotics lab has their own data format: Open-X uses RLDS, HuggingFace uses LeRobot, Diffusion Policy uses Zarr, robomimic uses HDF5, real-world ROS2 / teleop pipelines use MCAP. Want to train Octo on your ALOHA data? Write a converter. Want to use LeRobot on Open-X datasets? Write another. +
-Forge uses a hub-and-spoke architecture — one intermediate representation, O(n) format support: +**Data engine**  ·  [Catalog](#the-data-engine) · [Ingest & query](#1-ingest--query) · [Search](#2-search-semantically) · [Dedup & curate](#3-dedup--curate) · [Studio](#4-forge-studio) -``` -Any Reader → Episode/Frame → Any Writer -``` +
-Add a reader, get all writers for free. Add a writer, get all readers for free. No N×M conversion logic. See [docs/architecture.md](docs/architecture.md) for details. +**Reference**  ·  [Cloud storage](#cloud-storage-s3--gcs) · [Registry](#dataset-registry) · [Formats](#supported-formats) · [Command cheatsheet](#command-cheatsheet) · [Roadmap](#roadmap) -## Try it in 60 seconds (no install) +
-[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/arpitg1304/forge/blob/main/notebooks/forge_quickstart.ipynb)   **→** Pick a public LeRobot dataset, score every episode on 8 quality metrics, drill into the worst demos. No GPU. No auth. ~60 seconds wall-clock. +--- -## Quick Start +## Install ```bash pip install forge-robotics # base CLI + LeRobot v3 read/write -pip install "forge-robotics[mcap]" # add MCAP read/write -pip install "forge-robotics[rlds,lerobot]" # pick the formats you need -pip install "forge-robotics[s3]" # read from Amazon S3 (s3://) -pip install "forge-robotics[gcs]" # read from Google Cloud Storage (gs://) pip install "forge-robotics[all]" # everything ``` -That gives you the `forge` CLI: -```bash -forge inspect path/to/dataset -forge convert path/to/dataset ./out --format lerobot-v3 -forge visualize path/to/dataset -``` +Pick only the extras you need: -### Develop from source +| Extra | Adds | Extra | Adds | +|---|---|---|---| +| `[rlds]` | RLDS / Open-X (TensorFlow) | `[s3]` | Read from Amazon S3 (`s3://`) | +| `[lerobot]` | LeRobot v2/v3 (Parquet) | `[gcs]` | Read from Google Cloud (`gs://`) | +| `[mcap]` | MCAP (ROS2 + Foxglove) | `[catalog]` | The catalog (DuckDB) | +| `[hdf5]` | HDF5 (ALOHA, robomimic) | `[embed]` | Semantic search (SigLIP) | +| `[zarr]` | Zarr (Diffusion Policy, UMI) | `[video]` | Video quality + Studio thumbnails | +| `[rosbag]` | ROS1/ROS2 bags | `[rerun]` | Rerun 3D viewer | -```bash -git clone https://github.com/arpitg1304/forge.git -cd forge -pip install -e ".[all,dev]" -``` +> **Heads up:** the published PyPI build is the stable **format toolkit**. The **data-engine** features — catalog, semantic search, dedup/curation, and Forge Studio — are newer than the last release tag. To use those today, install from source: +> +> ```bash +> git clone https://github.com/arpitg1304/forge.git && cd forge +> pip install -e ".[all]" +> ``` -### RoboDM Support (Optional) +
+RoboDM (optional) -RoboDM requires manual installation from GitHub (PyPI version has a codec bug): +RoboDM's `.vla` format (up to 70× compression) needs a manual install — the PyPI build has a codec bug: ```bash git clone https://github.com/BerkeleyAutomation/robodm.git pip install -e robodm ``` +
-### Usage +## Try it in 60 seconds -```bash -# See what's in a dataset -forge inspect /path/to/dataset +[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/arpitg1304/forge/blob/main/notebooks/forge_quickstart.ipynb) — pick a public LeRobot dataset, score every episode on 8 quality metrics, drill into the worst demos. No GPU, no auth, ~60 seconds. -# Convert it -forge convert /path/to/rlds ./output --format lerobot-v3 -forge convert hf://arpitg1304/stack_lego ./stack_lego_rlds --format rlds --workers 4 --visualize -forge convert hf://lerobot/pusht ./pusht_robodm --format robodm -``` - -Works with HuggingFace Hub too: +Or locally: ```bash -forge inspect hf://lerobot/pusht -forge convert hf://lerobot/pusht ./output --format lerobot-v3 +forge inspect hf://lerobot/pusht # what's in it? +forge convert hf://lerobot/pusht ./out --format lerobot-v3 # convert it +forge quality hf://lerobot/pusht # score every episode ``` -### Cloud storage (S3 & GCS) +Forge speaks a hub-and-spoke architecture: **any reader → `Episode`/`Frame` → any writer**. Add a reader, get every writer for free — no N×M conversion logic. See [docs/architecture.md](docs/architecture.md). -Every command that takes a dataset path also accepts `s3://` and `gs://` URIs, -in addition to local paths and `hf://` URLs: +--- -```bash -pip install "forge-robotics[s3]" # or [gcs] for Google Cloud Storage +# The format toolkit -forge inspect s3://my-bucket/datasets/run_0413 -forge convert gs://lab-data/rosbags ./out --format lerobot-v3 -forge quality s3://my-bucket/datasets/droid --report report.html -``` - -Cloud datasets are downloaded to a temporary directory on first access and -cleaned up automatically when the command finishes. This keeps every format -(including video, HDF5, and rosbag, which need random file access) working -exactly as it does locally. - -**Authentication** uses each provider's standard credential chain — Forge never -handles credentials itself: +Per-dataset operations. Every command takes a local path, a registry id (`droid`), an `hf://` URL, or an `s3://` / `gs://` URI. -- **S3** — `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` env vars, `~/.aws/config` - profiles (`AWS_PROFILE`), or the instance/EKS IAM role. See the - [AWS credentials docs](https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html). -- **GCS** — Application Default Credentials: `gcloud auth application-default login`, - a service-account key via `GOOGLE_APPLICATION_CREDENTIALS`, or the attached - service account on GCP. See the - [GCP ADC docs](https://cloud.google.com/docs/authentication/application-default-credentials). +## Convert & interop -> **Writing** outputs directly to `s3://` / `gs://` is not supported yet — write -> to a local directory and upload it afterwards (`aws s3 cp --recursive`, -> `gcloud storage cp --recursive`). - -See [forge/io/README.md](forge/io/README.md) for the Python API and a guide to -diagnosing cloud bucket connectivity issues. - -## Common conversions +```bash +forge inspect ./dataset # structure, schema, cameras +forge convert ./rlds_dataset ./out --format lerobot-v3 # convert between formats +forge convert ./data.zarr ./out --format lerobot-v3 --visualize +``` | You have | You want | One command | |---|---|---| -| MCAP recording from ROS2 / teleop | LeRobot v3 for HuggingFace | `forge convert teleop.mcap ./out --format lerobot-v3` | -| RLDS from Open-X Embodiment | LeRobot for finetuning | `forge convert hf://openvla/modified_libero_rlds ./out --format lerobot-v3` | +| MCAP from ROS2 / teleop | LeRobot v3 for HuggingFace | `forge convert teleop.mcap ./out --format lerobot-v3` | +| RLDS from Open-X | LeRobot for finetuning | `forge convert hf://openvla/modified_libero_rlds ./out -f lerobot-v3` | | HDF5 from ALOHA / robomimic | MCAP for Foxglove playback | `forge convert aloha.hdf5 ./out --format mcap` | | Zarr from Diffusion Policy | LeRobot v3 | `forge convert pusht.zarr ./out --format lerobot-v3` | -| Any supported format | Quality scores per episode | `forge quality ./dataset` | -| Any supported format | Video quality (blur, motion, cuts) | `forge quality ./dataset --video --video-level motion` | -| Any supported format | Lint for hygiene defects | `forge lint ./dataset` | -| Any supported format | Filter out bad demos | `forge filter ./dataset ./clean --min-quality 6.0` | -| Any supported format | Remove near-duplicate episodes | `forge dedup ./dataset ./deduped` | -| Continuous actions | Discrete action tokens for VLA training | `forge tokenize write ./dataset ./tokenized --strategy openvla-bins` | -## Python API +For complex conversions, generate a YAML config: `forge inspect ds/ --generate-config config.yaml`, then `forge convert ds/ out/ --config config.yaml`. See [docs/configuration.md](docs/configuration.md). -```python -import forge - -# Inspect -info = forge.inspect("/path/to/dataset") -print(info.format, info.num_episodes, info.cameras) - -# Convert -forge.convert( - "/path/to/rlds", - "/path/to/output", - target_format="lerobot-v3" -) -``` - -## Quality Metrics +## Score & clean -Automated episode-level quality scoring from proprioception data alone — no video processing needed. +Quality — score each episode 0–10 from proprioception alone (no video needed), on 8 research-backed metrics. ```bash forge quality ./my_dataset -forge quality hf://lerobot/aloha_sim_cube --export report.json +forge quality ./my_dataset --video --video-level motion # also score camera streams ``` -Scores each episode 0-10 based on 8 research-backed metrics: +
+The 8 metrics -- **Smoothness (LDLJ)** — jerk-based smoothness from motor control literature (Hogan & Sternad, 2009) +- **Smoothness (LDLJ)** — jerk-based smoothness (Hogan & Sternad, 2009) - **Dead actions** — zero/constant action detection (Kim et al. "OpenVLA", 2024) - **Gripper chatter** — rapid open/close transitions (Sakr et al., 2024) -- **Static detection** — idle periods where the robot isn't moving (Liu et al. "SCIZOR", 2025) +- **Static detection** — idle periods (Liu et al. "SCIZOR", 2025) - **Timestamp regularity** — dropped frames and frequency jitter - **Action saturation** — time spent at hardware limits - **Action entropy** — diversity vs repetitiveness (Belkhale et al. "DemInf", 2025) -- **Path length** — wandering/hesitation in joint space - -See [forge/quality/README.md](forge/quality/README.md) for full metric details, paper references, and how to add new metrics. +- **Path length** — wandering / hesitation in joint space -### Video quality (opt-in) +`--video` adds pixel metrics (blur, exposure, frozen frames) and, at `--video-level motion`, optical-flow motion, camera-vs-scene split, and shot-cut detection. Needs `[video]`. Details: [quality](forge/quality/README.md) · [video quality](forge/quality/video/README.md). +
-Proprio scoring is the default fast lane; pass `--video` to also score the camera streams. **Tier 0** (`pixel`) adds sharpness/blur, exposure, frozen-frame, and colorfulness; **Tier 1** (`motion`) adds optical-flow motion magnitude, smoothness, camera-vs-scene split, and shot-cut detection. Add `--workers N` to parallelize. +Lint — check dataset *hygiene* (missing task strings, ambiguous cameras, low-res / single-view, missing action fields) against HuggingFace's LeRobot guidelines. Exits non-zero, so it drops into CI. ```bash -forge quality ./my_dataset --video # Tier 0 (pixel) -forge quality ./my_dataset --video --video-level motion # Tier 1 (optical flow) -forge filter ./my_dataset ./clean --min-sharpness 80 --min-motion 0.1 --exclude-flags cut_detected +forge lint ./my_dataset --strict # fail on warnings too ``` -Requires the `[video]` extra (`pip install forge-robotics[video]`). See [forge/quality/video/README.md](forge/quality/video/README.md). +Filter — drop bad episodes by quality, flags, or ids → a new dataset. -## Episode Filtering +```bash +forge filter ./my_dataset ./clean --min-quality 6.0 --exclude-flags jerky,mostly_static +``` -Filter datasets by quality score, flags, or episode IDs. Supports dry-run previews and pre-computed quality reports. +Dedup — remove near-duplicate episodes *within one dataset* by perceptual hashing of keyframes (numpy only, no model). ```bash -forge filter ./my_dataset --min-quality 6.0 # Dry-run preview -forge filter ./my_dataset ./filtered --min-quality 6.0 # Write filtered dataset -forge filter ./my_dataset ./filtered --exclude-flags jerky,mostly_static -forge filter ./my_dataset ./filtered --from-report report.json # Skip re-analysis +forge dedup ./my_dataset ./deduped --threshold 0.05 ``` -See [forge/filter/README.md](forge/filter/README.md) for full details. +Details: [lint](forge/lint/README.md) · [filter](forge/filter/README.md) · [dedup](forge/dedup/README.md). -## Dataset Linting +## Understand -Check a dataset against Hugging Face's published [LeRobot recording guidelines](https://huggingface.co/blog/lerobot-datasets) and flag hygiene defects *before* you spend GPU-hours training on it. Where `forge quality` scores trajectory *content* (smoothness, dead actions, chatter), `forge lint` checks *hygiene*: missing or placeholder task strings, ambiguous camera naming, low-resolution or single-view setups, and missing action fields. +Segment — split episodes into phases (reach / grasp / place) via PELT changepoint detection on proprio. ```bash -forge lint ./my_dataset -forge lint hf://lerobot/pusht --export lint.json -forge lint ./my_dataset --strict # fail on warnings too, not just errors +forge segment hf://lerobot/droid_100 --export segments.json --plot timeline.png ``` -Runs against the reader's inspected metadata — no video decode, no full episode scan. Exits non-zero on any error (or any warning under `--strict`), so it drops straight into CI. +Tokenize — turn continuous actions into discrete tokens for VLA training; benchmark strategies on *your* data. -See [forge/lint/README.md](forge/lint/README.md) for the full check list and thresholds. +```bash +forge tokenize compare ./my_dataset --sample 20 # recon error / vocab util +forge tokenize write ./my_dataset ./tokenized --strategy openvla-bins +``` -## Deduplication +Built-ins: `uniform-bins` (RT-1), `openvla-bins`, `quantile-bins`, `mu-law`. -Find and remove near-duplicate episodes (exact copies, re-encodes, near-identical takes) by perceptual hashing of per-camera keyframes — numpy only, no model. +Visualize — three backends: browser (default), matplotlib, and [Rerun](https://rerun.io) (cameras + time-series on one timeline). ```bash -forge dedup ./my_dataset # Dry-run: report duplicate clusters -forge dedup ./my_dataset ./deduped --threshold 0.05 # Write deduplicated dataset -forge dedup ./my_dataset ./deduped --method dhash # phash (default) | dhash | ahash +forge visualize pusht # web (no install) +forge visualize pusht --backend rerun --segment ``` -See [forge/dedup/README.md](forge/dedup/README.md) for the algorithm and tuning. +![Rerun viewer showing camera stream alongside action and state time series](docs/assets/rerun_viz.png) + +Details: [segment](forge/segment/README.md) · [tokenize](forge/tokenize/README.md). -## The catalog +--- -Forge is a per-dataset tool by default. The **catalog** turns it into a *system of record*: an append-only set of Parquet tables that registers every episode you ingest and annotates it with quality scores, all queryable with SQL. It's zero-server (just Parquet + embedded DuckDB), works on a local directory or an `s3://` / `gs://` bucket, and is readable by pandas/Polars/Spark without Forge. +# The data engine + +The toolkit works one dataset at a time. The **catalog** turns Forge into a *system of record*: an append-only set of Parquet tables that registers every episode you ingest and annotates it with quality, embeddings, and curation decisions — all queryable with SQL. It's **zero-server** (Parquet + embedded DuckDB), lives on a **local dir or an `s3://` / `gs://` bucket**, and is readable by pandas / Polars / Spark without Forge. + +``` +ingest ──▶ query ──▶ embed ──▶ search ──▶ dedup ──▶ curate ──▶ Studio + (snapshot → soon) +``` ```bash -pip install "forge-robotics[catalog]" +pip install "forge-robotics[catalog]" # + [embed] for search, [video] for Studio thumbnails +``` -# 1. Create a catalog (local dir or cloud bucket) -forge catalog init ./forge-catalog +## 1. Ingest & query -# 2. Ingest datasets — registers + quality-scores each episode. -# Re-running is a no-op (episodes are skipped by content hash). -forge ingest ./my_dataset --catalog ./forge-catalog -forge ingest s3://lab-bucket/raw/2026-07-18/ -c ./forge-catalog +```bash +forge catalog init ./forge-catalog # local dir or cloud bucket +forge ingest ./my_dataset -c ./forge-catalog # register + quality-score each episode +forge ingest s3://lab-bucket/raw/2026-07-18/ -c ./forge-catalog # re-runs are a no-op (content hash) -# 3. Query with SQL (views: episodes, quality_scores, v_latest_quality) forge query "SELECT task, count(*) FROM episodes GROUP BY task" -c ./forge-catalog -forge query "SELECT e.language_instruction, q.overall_score - FROM episodes e JOIN v_latest_quality q USING(episode_id) - ORDER BY q.overall_score DESC LIMIT 10" -c ./forge-catalog --format json - -# 4. Summary stats -forge catalog stats --catalog ./forge-catalog +forge catalog stats -c ./forge-catalog ``` -Python API: +Ingestion reuses the same readers as `forge inspect` and the same scorer as `forge quality`, so the catalog stays consistent with the toolkit. Writes go through pyarrow; reads through DuckDB (views: `episodes`, `quality_scores`, `v_latest_quality`). ```python from forge.catalog import Catalog @@ -287,192 +227,109 @@ df = cat.sql("SELECT robot, avg(overall_score) FROM episodes " "JOIN v_latest_quality USING(episode_id) GROUP BY robot").to_pandas() ``` -Ingestion reuses Forge's existing readers (the metadata behind `forge inspect`) and scorer (the engine behind `forge quality`), so the catalog stays consistent with the rest of the toolkit. Writes go through pyarrow; reads through DuckDB; nothing else touches catalog files. See [forge/catalog/README.md](forge/catalog/README.md) for the architecture, storage layout, and commit protocol. - -### Semantic search +## 2. Search semantically -Embed the episodes in a catalog, then search them by natural language — "find me the regrasp-after-a-failed-pick episodes" instead of scrolling folders. Uses [SigLIP](https://huggingface.co/google/siglip-so400m-patch14-384) (a shared image–text model), so text queries match episode *video*, not just metadata. +Embed episodes with [SigLIP](https://huggingface.co/google/siglip-so400m-patch14-384) (a shared image–text model), then search by natural language — text queries match episode *video*, not just metadata. ```bash pip install "forge-robotics[embed]" -# Embed every episode (vision per camera + instruction text). GPU auto-detected -# (CUDA → Apple MPS → CPU); re-running is a no-op. -forge embed --catalog ./forge-catalog - -# Search by text … +forge embed -c ./forge-catalog # GPU auto: CUDA → Apple MPS → CPU forge search "picks up the red cup" -c ./forge-catalog --top 10 -# … or find visually-similar episodes to one you already like -forge search --like -c ./forge-catalog +forge search --like -c ./forge-catalog # visually-similar episodes ``` -Vectors are versioned per model (`model_id = siglip-so400m@`) and stored in the same append-only catalog. Brute-force cosine in DuckDB is sub-second at lab scale. See [forge/embed/README.md](forge/embed/README.md) for models, device selection, and reproducibility. +Vectors are versioned per model (`siglip-so400m@`) and stored in the catalog. Details: [forge/embed/README.md](forge/embed/README.md). -### Dedup & curation +## 3. Dedup & curate -Use the embeddings to find near-duplicate episodes (re-encodes, near-identical retakes) and curate a clean, labeled training set. Near-dup pairs are recorded as **facts** (`dedup_edges`); which episode wins is decided at curation time by **policy**. +Find near-duplicate episodes across the whole corpus (cosine over embeddings), then curate a clean, labeled training set. Near-dup pairs are recorded as **facts**; which episode wins is decided by **policy** at curation time. ```bash -# Find near-duplicate pairs (cosine over episode embeddings) -forge catalog dedup -c ./forge-catalog --threshold 0.97 +forge catalog dedup -c ./forge-catalog --threshold 0.97 # store near-dup pairs (dedup_edges) -# Approve a high-quality selection, dropping dedup losers by policy forge curate -c ./forge-catalog \ --where "overall_score > 6 AND task = 'pick_place'" \ --dedup 0.97 --dedup-policy keep-higher-quality --label approved ``` -Curation is an append-log (`curation_labels`, latest-row-wins); nothing is ever deleted. Policies: `keep-higher-quality`, `keep-longer`, `keep-first`. +Curation is an append-log (`curation_labels`, latest-wins) — nothing is deleted. Policies: `keep-higher-quality`, `keep-longer`, `keep-first`. -### Forge Studio +## 4. Forge Studio -Generate a self-contained, themed HTML app to explore the catalog visually — Overview, Corpus (with thumbnails + quality rings), Dedup review (keep/reject pairs), and a Snapshot preview: +A self-contained, themed HTML app to explore the catalog — **Overview · Corpus** (thumbnails + quality rings) **· Dedup review** (keep/reject pairs) **· Snapshot**. One shareable file, no server; real data and video thumbnails embedded. ```bash forge studio -c ./forge-catalog -o studio.html && open studio.html ``` -Everything is embedded (real data + video thumbnails as data URIs) — one shareable file, no server. See [forge/catalog/README.md](forge/catalog/README.md). +Details, storage layout, and commit protocol: [forge/catalog/README.md](forge/catalog/README.md). A ready-to-explore [example catalog](forge/catalog/catalog_example_droid_100/) (droid_100) ships in the repo. -## Dataset Registry +--- -A curated catalog of 23+ prominent robotics datasets — browse, search, and download by name instead of memorizing URIs. **[Browse the registry online](https://arpitg1304.github.io/forge/registry.html)** +# Reference -```bash -# Browse all datasets -forge registry list +## Cloud storage (S3 & GCS) -# Open an interactive HTML browser with filtering -forge registry list --html - -# Filter by format, embodiment, or tags -forge registry list --format rlds --embodiment franka -forge registry list --tag manipulation --demo - -# Get detailed info on a dataset -forge registry info droid - -# Search across names, tags, embodiments, and task types -forge registry search "franka manipulation" - -# Validate the registry (for contributors) -forge registry validate -``` - -### Registry ID Resolution - -Use dataset IDs directly in any command — no need for full paths or URIs: - -```bash -forge inspect droid # resolves to hf://lerobot/droid -forge quality pusht # resolves to hf://lerobot/pusht -forge convert droid ./output --format lerobot-v3 -``` - -### Quick Start with `forge demo` - -Download a small demo dataset, inspect it, and run quality scoring — all in one command: - -```bash -forge demo # uses pusht by default -forge demo aloha_sim_cube # or pick any demo-suitable dataset -``` - -See [forge/registry/CONTRIBUTING.md](forge/registry/CONTRIBUTING.md) for how to add new datasets to the registry. - -## Episode Segmentation - -Automatic episode segmentation via PELT changepoint detection on proprioception signals. Splits episodes into contiguous phases (sub-skills, regime changes, idle periods) without video processing. - -```bash -forge segment ./my_dataset -forge segment hf://lerobot/droid_100 --export segments.json --plot timeline.png -forge segment ./my_dataset --signal action --penalty bic --cost-model rbf -forge segment ./my_dataset --sample 20 -``` - -Detects where the statistical properties of the proprio signal change abruptly — e.g., transitions between reaching, grasping, and placing phases. Configurable cost models (`rbf`, `l2`, `l1`), penalty methods (`bic`, `aic`, or numeric), and signal selection (`observation.state`, `action`, `qpos`). - -See [forge/segment/README.md](forge/segment/README.md) for full details. - -## Action Tokenization - -Turn continuous action vectors into discrete tokens (and back) for VLA / robot-learning models, which predict discrete action tokens rather than continuous vectors. Proven strategies ship in-box; a comparator benchmarks them on *your* dataset so you don't have to guess. +Every command that takes a dataset or catalog path also accepts `s3://` and `gs://` URIs. Cloud datasets are downloaded to a temp dir on first access and cleaned up automatically, so every format works exactly as it does locally. ```bash -forge tokenize list # registered strategies -forge tokenize compare ./my_dataset --sample 20 --export report.json # benchmark recon error / vocab util -forge tokenize fit ./my_dataset --strategy openvla-bins --out tok.json -forge tokenize write ./my_dataset ./tokenized --strategy openvla-bins # LeRobot v3 + action_tokens column +pip install "forge-robotics[s3]" # or [gcs] +forge inspect s3://my-bucket/datasets/run_0413 +forge convert gs://lab-data/rosbags ./out --format lerobot-v3 ``` -Built-in strategies: `uniform-bins` (RT-1), `openvla-bins` (OpenVLA), `quantile-bins`, and `mu-law` — all per-step and numpy-only. `write` saves the fitted tokenizer to `meta/action_tokenizer.json` for inference-time detokenization. Add your own strategy with a one-line registry decorator. - -See [forge/tokenize/README.md](forge/tokenize/README.md) for full details and the extension API. +Auth uses each provider's standard credential chain (AWS env vars / profiles / IAM roles; GCP Application Default Credentials) — Forge never handles credentials itself. Catalogs can be read *and written* in the cloud; per-format conversion **outputs** are local-only for now. Full guide + connectivity troubleshooting: [forge/io/README.md](forge/io/README.md). -## Visualization +## Dataset registry -Forge ships three visualization backends selectable with `--backend`: +A curated catalog of 23+ prominent robotics datasets — browse and use by name instead of memorizing URIs. **[Browse online](https://arpitg1304.github.io/forge/registry.html)** ```bash -forge visualize pusht # web (default) — browser-based, no install -forge visualize pusht --backend matplotlib # matplotlib — sliders, comparison mode -forge visualize pusht --backend rerun # Rerun — cameras + time-series on one timeline -forge visualize pusht --backend rerun --segment # with PELT phase labels -forge visualize pusht --backend rerun --samples 3 # stream multiple episodes +forge registry list --format rlds --embodiment franka # filter +forge registry search "franka manipulation" +forge inspect droid # ids work in any command +forge demo # download + inspect + score a demo ``` -The **Rerun backend** logs each frame's camera images, per-dimension action and state scalars, and segment labels into the [Rerun](https://rerun.io) viewer — all aligned on a shared `frame` timeline. - -![Rerun viewer showing camera stream alongside action and state time series](docs/assets/rerun_viz.png) +Add datasets via [forge/registry/CONTRIBUTING.md](forge/registry/CONTRIBUTING.md). -Install the Rerun extra to use it: - -```bash -pip install "forge-robotics[rerun]" -``` +## Supported formats -## CLI Reference +| Format | Read | Write | Notes | +|--------|:----:|:-----:|-------| +| RLDS | ✓ | ✓ | Open-X, TensorFlow Datasets | +| LeRobot v2/v3 | ✓ | ✓ | HuggingFace, Parquet + MP4 | +| GR00T | ✓ | – | NVIDIA Isaac, LeRobot v2 + embodiment metadata | +| RoboDM | ✓ | ✓ | Berkeley `.vla`, up to 70× compression (manual install) | +| Zarr | ✓ | – | Diffusion Policy, UMI | +| HDF5 | ✓ | – | robomimic, ACT/ALOHA | +| MCAP | ✓ | ✓ | ROS2 CDR + Foxglove Protobuf, no ROS install required | +| Rosbag | ✓ | – | ROS1 `.bag`, ROS2 SQLite3 | -See [docs/cli.md](docs/cli.md) for the full command reference including: +Which models use which format: [docs/model_formats.md](docs/model_formats.md) · format specs: [docs/format_reference.md](docs/format_reference.md). -- `forge inspect` - Dataset inspection and schema analysis -- `forge convert` - Format conversion with camera mapping -- `forge visualize` - Interactive dataset viewer (backends: `web`, `matplotlib`, `rerun`) -- `forge quality` - Episode-level quality scoring ([details](forge/quality/README.md)) -- `forge filter` - Quality-based episode filtering ([details](forge/filter/README.md)) -- `forge registry` - Browse and search the dataset registry -- `forge demo` - Quick-start with a demo dataset -- `forge segment` - Episode segmentation via changepoint detection ([details](forge/segment/README.md)) -- `forge stats` - Compute dataset statistics -- `forge export-video` - Extract camera videos as MP4 -- `forge hub` - Search and download from HuggingFace +## Command cheatsheet -## Configuration +**Toolkit** (per dataset): `inspect` · `convert` · `quality` · `lint` · `filter` · `dedup` · `segment` · `tokenize` · `visualize` · `stats` · `export-video` -For complex conversions, use a YAML config: +**Data engine** (per catalog): `catalog init` · `ingest` · `query` · `catalog stats` · `embed` · `search` · `catalog dedup` · `curate` · `studio` -```bash -forge inspect my_dataset/ --generate-config config.yaml -forge convert my_dataset/ output/ --config config.yaml -``` +**Discovery**: `hub` · `local` · `registry` · `demo` · `formats` · `version` -See [docs/configuration.md](docs/configuration.md) for details. +Full reference: [docs/cli.md](docs/cli.md). Run `forge --help` for any command. ## Roadmap -Planned features (contributions welcome!): - -- [ ] **Dataset merging** - Combine multiple datasets into one (`forge merge ds1/ ds2/ --output combined/`) -- [ ] **Train/val/test splitting** - Split datasets with stratification (`--split 80/10/10`) -- [x] **Dataset registry** - Curated catalog of 23+ robotics datasets with CLI browser and HTML viewer -- [x] **MCAP first-class support** - Read + write, ROS2 CDR + Foxglove Protobuf, no ROS install required -- [ ] **Streaming reads** - Process HuggingFace datasets without full download -- [x] **Episode filtering** - Filter by quality score, flags, or episode IDs (`forge filter --min-quality 6.0`) -- [ ] **Depth/point cloud support** - Preserve depth streams from RLDS/Open-X -- [ ] **GR00T writer** - Write to NVIDIA Isaac GR00T training format (read support complete) -- [ ] **Distributed conversion** - Scale to 100K+ episode datasets across nodes -- [ ] **Conversion verification** - Automated diff between source and converted data +- [x] Dataset registry — curated catalog of 23+ datasets with CLI + HTML browser +- [x] MCAP first-class support — read + write, ROS2 CDR + Foxglove, no ROS install +- [x] Episode filtering, quality scoring, linting, segmentation, tokenization +- [x] Cloud storage — `s3://` / `gs://` on every command +- [x] The catalog — ingest, SQL query, semantic search, dedup, curation, Studio +- [ ] **Snapshots + export** — freeze a curated selection → LeRobot/RLDS for training +- [ ] **Streaming reads** — process cloud datasets without a full download +- [ ] **Dataset merging & splitting** — combine datasets; stratified train/val/test +- [ ] **Depth / point-cloud support** · **GR00T writer** · **distributed conversion** ## Development @@ -482,6 +339,8 @@ make install-dev make test ``` +Contributions welcome. See [docs/architecture.md](docs/architecture.md) for the design. + ## License MIT diff --git a/docs/index.html b/docs/index.html index a469fa3..84c3dca 100644 --- a/docs/index.html +++ b/docs/index.html @@ -3,8 +3,8 @@ -Forge - Robotics Data Toolkit - +Forge - Robotics Data Toolkit & Data Engine +