Skip to content

Mpc stage1 - #1220

Open
TamasTran wants to merge 46 commits into
mujocolab:mainfrom
TamasTran:mpc-stage1
Open

TamasTran wants to merge 46 commits into
mujocolab:mainfrom
TamasTran:mpc-stage1

Conversation

@TamasTran

Copy link
Copy Markdown

No description provided.

TamasTran and others added 30 commits September 25, 2026 15:22
vtrace() called rhos.clamp(max=None), which raises in torch when
importance_weight_clip_max=None (the "no truncation" ablation). Skip the
clamp in that case.

Also make the replay buffer and its test pass ty/pyright: assign TensorDict
slices instead of copy_(), cast observation groups when allocating, and
narrow the Optional observations in the test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…switch

MpopiPpo subclasses RSL-RL's PPO and is selected through algorithm.class_name, so rsl-rl-lib itself is untouched. It records raw rewards and time-outs during the rollout (before PPO folds the stale time-out bootstrap into the stored rewards), asks MPOPI for importance-corrected replay samples at the start of each update, and optimizes fresh plus replay samples with PPO's clipped objective weighted by clip(pi_old / mu). There is still a single optimizer. update() mirrors rsl-rl-lib 5.5.1 because the upstream loss has no hook for per-sample weights; with no replay available it follows the upstream code path exactly, and a test checks that parameters are bit-identical to plain PPO.

The new RslRlPpoAlgorithmCfg.mpopi field defaults to mode="ppo", in which case MjlabOnPolicyRunner removes it and constructs upstream PPO with the same arguments as before. The naive_replay_ppo and mpopi_ppo modes are selectable from the CLI with --agent.algorithm.mpopi.mode. KL and clip fraction are now logged in replay modes along with the MPOPI diagnostics.

A small pure-torch point-mass VecEnv is included for fast CPU tests and the upcoming toy benchmark.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/benchmarks/mpopi_toy_benchmark.py runs PPO, naive replay + PPO and MPOPI + PPO (with and without truncation) on the pure-torch point-mass env through the real runner mode switch, evaluates the deterministic policy after every update, and reports AUC, final return and Welch t-tests across seeds.

docs/mpopi_toy_results.md records an exploratory run, a confirmation run on fresh seeds and replay-ratio / buffer-size ablations. On this toy task MPOPI + PPO never collapsed (0/80 runs) while naive replay collapsed in 11/70, and MPOPI's final return beat PPO by about 7% (confirmed, p=5.5e-5). No significant sample-efficiency (AUC) gain over PPO was found.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rename scripts/benchmarks/mpopi_toy_benchmark.py to mpopi_benchmark.py and add a --task option that runs a registered mjlab task with its own PPO settings. Task runs are scored by a deterministic-policy evaluation on a separate play env with a fixed seed; the Python, NumPy and torch RNG states are saved and restored around eval-env construction and every evaluation (mjlab env construction reseeds all three globally), and training curves were verified to be identical with and without per-iteration evaluation. Adds a compute-matched A_ppo_bigmb arm and Mann-Whitney tests. The toy benchmark reproduces its Phase 8 numbers exactly.

docs/mpopi_cartpole_results.md records a pre-registered study on Mjlab-Cartpole-Balance (256 envs, 5 seeds per arm). Neither hypothesis was supported: MPOPI + PPO did not learn faster than PPO or a compute-matched PPO, and in an exploratory paired comparison it was slower than both PPO and naive replay in 5/5 seeds. Removing truncation recovered part of the gap. Naive replay did not collapse on this task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
notebooks/mpopi_colab.ipynb installs the environment with uv on a Colab GPU runtime, runs the MPOPI unit tests and a short GPU smoke test, runs the A/B/C benchmark on a task with plots and paired tests, optionally trains a robot task in ppo / naive_replay_ppo / mpopi_ppo mode, and shows TensorBoard. The code can come from the GitHub branch or an uploaded git archive.

scripts/benchmarks/mpopi_benchmark.py gains --device so it can run on CUDA, and the RNG guard around evaluation now also saves and restores the CUDA generators. CPU results are unchanged (the Phase 8 toy numbers still reproduce exactly). The CUDA path has not been run yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The code step only checked that the mpopi package existed, so an older push of the branch (without MpopiPpo, the algorithm tests or the benchmark script) got through and pytest later failed with a file-not-found usage error. It now checks for the files the later steps need and says to push the branch or use the zip upload.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The benchmark now prints a progress line every --progress-every iterations with the latest eval score, the training reward and, for replay arms, KL, clip fraction, ESS, mean importance weight and KL(mu || pi_old). Results are unchanged; the Phase 8 toy numbers still reproduce exactly.

In the Colab notebook, robot training (section 7) now runs in the background with one console log per mode, so other cells can run meanwhile. New cells show the latest log block and a mean-reward curve parsed from those logs, wait for or stop the background run, and record a video from the newest checkpoint with play (usable mid-training). TensorBoard auto-reloads every 30 s. The log parser and the play video flow were checked locally against a real Cartpole training run; play needs --video True because mjlab's CLI takes explicit booleans.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mjlab.mpc.SamplingMpc plans for N real envs on a planning copy of the task with N x K worlds (no auto-reset, no terminations). Each control step copies the real envs' simulation state into their K planning worlds, rolls out K perturbed action sequences over the horizon with the task's own reward manager, and combines them with MPPI weights whose temperature is chosen per env to hit a target effective sample size. iterations > 1 adapts a per-dimension std between batches (MPOPI). Actions are in the policy action space, and the planner uses its own generator and restores the global RNG around planning-env construction.

For Cartpole, copying qpos, qvel, act, qacc_warmstart and ctrl reproduces the real env bit for bit, which the tests check. scripts/mpc/eval_mpc.py evaluates the planner as a controller: on 16 play envs x 200 steps MPPI scores 0.973 normalized (criterion >= 0.9 set beforehand), versus 0.695 for zero action and 0.300 for random actions. This is stage 1 of the MPC -> PPO plan; the planner is not connected to PPO yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MpcCollector drives extra envs with the sampling MPC, executes Gaussian noise around the MPC action and records the exact behavior density. In mode mpc_ppo, MpopiPpo adds these segments to PPO's updates through the existing MPOPI correction, plus a behavior-cloning term toward the MPC action that decays to zero. SamplingMpc now takes an env config instead of a task id. On Cartpole (5 seeds) the gain over PPO comes from behavior cloning; see docs/mpc_ppo_results.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With an imperfect MPC teacher (0.77 of the maximum as a controller), MPC data speeds up early learning but not the AUC over 200 iterations. Behavior cloning alone matches the full method, the importance correction mainly prevents the harm of uncorrected MPC samples, and corrected samples without behavior cloning do not help (5 seeds).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
At equal simulation budget MPOPI scores higher than MPPI as a controller, but as a teacher for PPO (K = 8, L = 4) it collected lower-reward data under execution noise and gave lower AUC on 4 of 5 seeds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ng-up recipe

MpcDataCfg gains execution_std=0 (clone the MPC action itself), driver="policy" (DAgger: the policy acts, the MPC labels its states) and bc_floor. Behavior-cloning-only updates now build their batch straight from the buffer, without importance weights, so data without a behavior density can be used.

On swing-up (5 seeds), BC with DAgger labels from a noise-free MPOPI teacher had the highest AUC and beat plain PPO on 4 of 5 seeds without the drop seen after BC fades. BC-floor runs crashed when the scalar actor std went negative. See docs/mpc_ppo_swingup_recipe.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runs PPO and the chosen DAgger recipe on swing-up for fresh seeds 500-509 on a Colab GPU, skipping finished runs after a disconnect, and tests the pre-registered hypothesis (DAgger AUC > PPO AUC, paired Wilcoxon, alpha 0.05).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a Colab GPU with fresh seeds 500-509, DAgger beat plain PPO on AUC in 7 of 10 seeds (+0.0058, Wilcoxon p = 0.13), so the pre-registered hypothesis was not confirmed. DAgger reached the 0.040 threshold on 9 of 10 seeds versus 4 of 10, but the CPU observation of no drop after behavior cloning ends did not replicate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
inject_fraction draws MPC samples uniformly from the buffer so that they make up a fixed fraction of every PPO batch for the whole run, following MPC-Injection (arXiv 2606.26392) but for PPO. Naive injection (correction=False) now accepts noise-free MPC data; the behavior KL metric is skipped when the behavior has no density.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runs PPO, MPC-DAgger and MPC-Injection (25% MPC samples per PPO batch) on swing-up for seeds 500-509, reusing the earlier DAgger results from the same Drive folder when present, and tests the pre-registered hypotheses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On swing-up, injecting noise-free MPC samples into PPO batches was worse than plain PPO in every finished run, and 10 of 15 runs crashed when the scalar actor std went negative, more often at higher injected fractions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With 10 seeds on a Colab GPU, MPC-Injection (25% of each PPO batch) did not beat plain PPO (3/8 seeds, p = 0.84) and was worse than MPC-DAgger (1/8 seeds, p = 0.023); 2 of 10 runs crashed with the negative-std error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the train CLI

Uses the standard train command with TensorBoard logging and checkpoints, runs the three methods in the background for several seeds, and compares training curves and play videos.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds a preset: the documented 4096 envs and 500 iterations (with 64 MPC envs x 32 steps so MPC-Injection still gets a 25% share), or the 64-env, 200-iteration setting of the reported experiments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The progress cell now reports the last iteration even before a mean reward exists (about iteration 31 on swing-up) and how recently each log changed, and training logs are unbuffered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The runner can raise the lower bound of a Gaussian actor's std in every mode, including plain PPO. The crashes previously described as a negative std were a NaN std: RSL-RL clamps the std to at least 1e-6, so the actor parameters had become NaN; the docs are corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The planner now copies commands and their timers, previous actions, stateful reward terms, sensor histories, entity data such as encoder bias, and domain-randomized model fields, not only the simulator state, and drops interval events and command resampling in the planning copy. On G1 velocity tracking the planner's per-term rewards now match the real env to about 1e-5; copying only the simulator state gave errors of order 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ation

num_knots samples the planner noise at a few knots and interpolates linearly, like MuJoCo MPC's spline plans; independent per-step noise is heavily penalized by action-rate terms on G1. scripts/mpc/eval_mpc_velocity.py drives a play env at fixed forward speeds with MPC or zero actions and reports the achieved speed, tracking error and falls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ript

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runs the pre-registered planner sweep of docs/mpc_g1_velocity.md at 0.5, 1.0 and 1.5 m/s on a GPU, applies the pass criterion, and shows videos.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…omparison notebook

Mjlab-Velocity-Flat-Unitree-G1-2k keeps the flat G1 task but trains for 2000 iterations and widens the forward command range from (-1, 1) to (-1, 1.5) m/s at iteration 500, instead of after 5000 iterations as in the 30000-iteration default. eval_mpc_velocity.py gains a policy controller that evaluates a checkpoint at fixed forward speeds. The notebook trains PPO, MPC-DAgger and MPC-Inject on this task with the same action std floor and measures the achieved speed at 0.5, 1.0 and 1.5 m/s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TamasTran and others added 14 commits September 29, 2026 15:34
…tebook

The new mpc_ppo option replay_own_rollouts keeps a second replay buffer of PPO's own past rollouts, corrected by MPOPI as in mpopi_ppo, while the MPC buffer only feeds the behavior-cloning loss. This allows DAgger and Replay-IS in one run. The G1 comparison notebook now trains PPO, Replay-IS, MPC-DAgger, Replay-IS + DAgger and MPC-Inject.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runs one training per T4 GPU in parallel, skips runs that cannot finish within the Kaggle session budget, resumes from a previous version's output attached as input, and packs the results into the notebook output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The new mpc_ppo option teacher_gap_every rolls out, from the same state and for the planner horizon, the plan the planner just optimized (open loop) and the policy's mean action (closed loop), and logs the return gap and the fraction of states where the plan is better. This tells on the actual task when the student overtakes the teacher, instead of assuming it from Cartpole.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tions

All runs train the full 2000 iterations. Besides PPO and the original Replay-IS + DAgger, configuration A halves the replay ratio and the MPC labeling envs, and configuration B uses 3072 main envs with half the MPC labeling envs. The notebook reports total time, seconds per iteration, iterations and minutes to each milestone, and the final tracking at 1.5 m/s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sults

Replay segments now go through the actor and critic in one batch instead of one forward pass per segment, and the MPC planner skips observation computation during its rollouts, where only rewards are used. Both produce the same values as before (checked against the previous implementation and by the existing planner test that compares rollout returns with full env steps).

MPOPI updates also log the wall time of each phase (time/mpc, time/replay, time/sgd, time/rest), and a Kaggle notebook reruns the original Replay-IS + DAgger configuration to measure the speed-up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The MPOPI replay correction, the sampling MPC teacher and the G1 2000-iteration task used to live inside the mjlab package and patched mjlab's runner, RL config and G1 task files. They now form their own package, src/mpopi_train, that uses mjlab as a library, so src/mjlab is identical to upstream again and merging upstream changes no longer touches research code.

The package adds MpopiRunnerCfg (mjlab's runner config with algorithm.mpopi), runners that mix the MPOPI setup into mjlab's runners instead of editing them, presets for the four compared methods, and one task per method (Mpopi-G1-2k-PPO, -Replay-IS, -DAgger, -Replay-IS-DAgger). The presets reproduce exactly the agent and env configs of the earlier Kaggle runs. New commands mpopi-train, mpopi-play and mpopi-eval wrap mjlab's train and play with these tasks registered, so a run is now `uv run mpopi-train Mpopi-G1-2k-Replay-IS-DAgger --agent.seed 1` instead of about twenty flags.

Research notes move to docs/mpopi, the Kaggle comparison notebook uses the new commands, and the older notebooks clone the tag mpopi-before-module so they keep running the code they were written for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The fixed-speed evaluation with 16 robots for 5 seconds was too noisy to compare methods: the same configuration gave tracking errors from 0.037 to 0.071 m/s at 1.5 m/s. A new Kaggle notebook finds the final checkpoints of earlier runs among its inputs and re-evaluates them with 64 robots over 10 seconds, reporting each run next to its old result and the mean and spread per method. The training notebook now evaluates the same way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The batched replay evaluation, the observation-free MPC rollouts and the per-phase timers saved about one minute per two-hour run, and the two runs made with them were the slowest and least accurate of six Replay-IS + DAgger runs. Their results are most likely run-to-run variance, but the saving does not justify the doubt, so the algorithm and MPC code go back to exactly what produced the earlier results (b3c8fe4), apart from import paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README keeps mjlab's structure and examples, and adds what the fork is for, how to train and evaluate the four compared methods, the main G1 results across all runs, the project layout and the notebooks. The guide is renamed to GUIDE.md and now evaluates with 64 robots over 10 seconds, explains how to read results and run-to-run variation, lists the Kaggle notebooks, and asks for one Kaggle notebook per experiment so earlier outputs stay reachable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The guide now has a full section on command-line parameters: the general syntax (pairs of numbers are written as -1.0,2.0), training settings, the target speed and its curriculum step (counted in simulation steps, iterations times 24), and the Replay-IS and DAgger settings with their defaults, with worked examples for a 2.0 m/s target and longer behavior cloning. The Kaggle training notebook gains EXTRA_ARGS and EVAL_SPEEDS so the same parameters can be set there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 7, 2026 09:55

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

TamasTran and others added 2 commits October 9, 2026 09:17
compare_with_policy fed the policy the planning env's cached observation from the end of the previous planning rollout instead of the copied state, so the first action of every teacher-gap comparison was wrong and the logged gap was biased toward the teacher. It now senses and computes the observation of the copied state; a new test checks the first observation the policy sees.

The MPC collector's envs now take the training env's step counter before each collection, so step-based curricula such as the command range follow training instead of staying at their first stage. A behavior-cloning floor with a finite max_age is rejected, since evicted labels would silently end the cloning. The Mpopi tasks are registered through the mjlab.tasks entry point, so worker processes of multi-GPU training and mjlab's own train and play find them. mpopi-eval loads the policy without building the MPC teacher and rejects settle_steps >= steps. The G1 curriculum takes the steps per iteration from the runner config, the benchmark imports at module level, long lines are wrapped, and the changelog describes the package.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… teacher

Replay-IS + DAgger is trained in three variants with two seeds each: the current schedule (labels until iteration 110, cloning weight reaches 0 at 150), labels until 300 with cloning until 340, and a permanent cloning weight of 0.1 on the last labels. Every variant measures the teacher-versus-policy gap with the fixed observation, and the notebook reports learning milestones, wall time, the cloning weight and loss, and the speed and error at 1.5 m/s. The base variant also checks that the current code reproduces the earlier runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants