仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
-
Updated
Oct 2, 2026 - Jupyter Notebook
仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
Robotics Data Toolkit | Convert between robotics dataset formats (RLDS, LeRobot v2/v3, Zarr, HDF5, Rosbag). Inspect, visualize, and analyze datasets. Works with HuggingFace Hub. Built for OpenVLA, Octo, LeRobot, and Diffusion Policy workflows.
PickAgent: OpenVLA-powered Pick and Place Agent | Gradio&Simulation | Vision Language Action Model
ROS2-native runtime, benchmark, and adapter hub for Vision-Language-Action models.
Train and deploy Vision-Language-Action models natively on Apple Silicon.
Independent VLA research notes: OpenVLA / π0 / Spirit paper reviews, LIBERO reproduction, XLeRobot integration. Transitioning from AD motion planning to embodied AI.
Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models
Using VLM-based visual question answering to perceive scenes and control robots in MuJoCo simulation.
FR3 robot in Gazebo integrated with 4-bit quantized OpenVLA and MoveIt
A full-stack Embodied AI simulation suite powered by Genesis World: featuring OpenVLA closed-loop evaluation, procedural scene generation, massively parallel GPU RL (Unitree Go2), and multi-physics coupling (PBD/SPH).
A RoboNix Skill for experience-memory retrieval and verified action reuse across OpenVLA and π0.
Safety monitors for learned robot policies under distribution shift. A benchmark for which monitor still works once the deployment distribution moves.
Do robot policies actually use their cameras? Reproducing ACT, OpenVLA, SmolVLA, Diffusion Policy, π0 and π0.5 in simulation, plus vision-dependence experiments (blackout, perturbation, trajectories, state-zeroing fine-tune).
Reproducible audit of SAFE VLA failure detection on OpenVLA/LIBERO — pooled AUC inflates +0.09 over macro within-task; functional CP catches only 7-9% at ≤5% FPR
Unofficial PyTorch reproduction for OpenVLA: An Open-Source Vision-Language-Action Model.
Simulated UR5e + MoveIt 2 cell for testing vision-language-action policies in ROS 2 Jazzy. Swap SmolVLA, OpenVLA-7B (4-bit) or GR00T N1.7 behind one ZeroMQ protocol — all runnable on a 6 GB GPU.
VLA manipulation ablations, held-out generalization, and an OpenVLA-7B feasibility study, all measured on a 6GB laptop GPU
π0.5 hears the instruction. What it hears does not decide what it does. Paired-prompt probing and causal intervention on OpenVLA and π0.5 (LIBERO), with an interactive explorer.
To associate your repository with the openvla topic, visit your repo's landing page and select "manage topics."