diff --git a/README.md b/README.md index 7325df4..4d0caae 100644 --- a/README.md +++ b/README.md @@ -7,275 +7,215 @@ ██║ ╚██████╔╝██║ ██║╚██████╔╝███████╗ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝ -
RLDS ═══╗ ╔═══► LeRobotHDF5 ═══╣ ╠═══► MCAPZarr ═══╬════⚙════╬═══► RoboDMMCAP ═══╝ ╚═══► RLDS
-Convert between robotics dataset formats with one command. Score demonstration quality with research-backed metrics. Lint datasets for hygiene defects before training. Segment episodes into sub-skills with changepoint detection.
+Forge is two things that share one core:
-| Format | Read | Write | Visualize | Notes |
-|--------|:----:|:-----:|:---------:|-------|
-| RLDS | ✓ | ✓ | ✓ | Open-X, TensorFlow Datasets |
-| LeRobot v2/v3 | ✓ | ✓ | ✓ | HuggingFace, Parquet + MP4 |
-| GR00T | ✓ | - | ✓ | NVIDIA Isaac, LeRobot v2 with embodiment metadata |
-| RoboDM | ✓ | ✓ | ✓ | Berkeley's .vla format, up to 70x compression* |
-| Zarr | ✓ | - | ✓ | Diffusion Policy, UMI |
-| HDF5 | ✓ | - | ✓ | robomimic, ACT/ALOHA |
-| MCAP | ✓ | ✓ | ✓ | ROS2 CDR + Foxglove Protobuf, no ROS install required |
-| Rosbag | ✓ | - | ✓ | ROS1 .bag, ROS2 SQLite3 |
+- **A format toolkit** — convert, inspect, score, lint, filter, segment, and visualize a single dataset across RLDS, LeRobot, HDF5, MCAP, Zarr, Rosbag, and more.
+- **A data engine** — register every episode you collect into an append-only **catalog**, then query it with SQL, search it by natural language, dedup and curate it, and explore it in a visual **Studio**.
-*\*RoboDM requires manual installation from GitHub (see below)*
+Everything works on local paths, `hf://` datasets, and `s3://` / `gs://` buckets.
-See [docs/model_formats.md](docs/model_formats.md) for which models (Octo, OpenVLA, ACT, Diffusion Policy, etc.) use which format. See [docs/format_reference.md](docs/format_reference.md) for detailed format specifications.
+| -## Why Forge? +**Toolkit** · [Install](#install) · [Convert](#convert--interop) · [Quality](#score--clean) · [Filter](#score--clean) · [Segment](#understand) · [Tokenize](#understand) · [Visualize](#understand) -Every robotics lab has their own data format: Open-X uses RLDS, HuggingFace uses LeRobot, Diffusion Policy uses Zarr, robomimic uses HDF5, real-world ROS2 / teleop pipelines use MCAP. Want to train Octo on your ALOHA data? Write a converter. Want to use LeRobot on Open-X datasets? Write another. + |
| -Forge uses a hub-and-spoke architecture — one intermediate representation, O(n) format support: +**Data engine** · [Catalog](#the-data-engine) · [Ingest & query](#1-ingest--query) · [Search](#2-search-semantically) · [Dedup & curate](#3-dedup--curate) · [Studio](#4-forge-studio) -``` -Any Reader → Episode/Frame → Any Writer -``` + |
| -Add a reader, get all writers for free. Add a writer, get all readers for free. No N×M conversion logic. See [docs/architecture.md](docs/architecture.md) for details. +**Reference** · [Cloud storage](#cloud-storage-s3--gcs) · [Registry](#dataset-registry) · [Formats](#supported-formats) · [Command cheatsheet](#command-cheatsheet) · [Roadmap](#roadmap) -## Try it in 60 seconds (no install) + |