Skip to content

feat: add deterministic evaluation runner - #2772

Merged
lvyufeng merged 1 commit into
masterfrom
feature/mindact-pr1-evaluation
Aug 31, 2026
Merged

lvyufeng merged 1 commit into
masterfrom
feature/mindact-pr1-evaluation

Conversation

@lvyufeng

@lvyufeng lvyufeng commented Aug 31, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds the first working piece of MindAct beyond the skeleton: a deterministic, dependency-free evaluation runner with immutable experiment provenance. No LeRobot, LIBERO, ACT or training-loop code is introduced here; this PR establishes the contract those integrations will plug into.

What's in it

Evaluation runner (src/mindact/evaluation/runner.py)

  • Episode N is seeded with config.seed + N, so a run is reproducible from its manifest alone.
  • Loop is reset(seed=...) -> predict(observation) -> step(action) until done or evaluation.max_steps; hitting the step limit sets truncated=True.
  • environment.close() always runs in a finally block, including when the policy raises.
  • Adapters are injected (policy instance + environment factory), so a real simulator can be connected without touching experiment bookkeeping.
  • Validates what adapters return: numeric finite reward, boolean done, mapping info, JSON-serializable metadata.

Write-once provenance

  • An existing run directory is never reused or overwritten; manifest.json and result files are created via a same-directory temp file, fsync, then os.link() for exclusive creation.
  • The manifest is written before the first rollout, so provenance survives a crash mid-run.
  • Aggregate metrics (evaluation/results.json) and per-episode records (evaluation/episodes.jsonl) are separate files, so completing an evaluation never mutates the manifest.
  • The manifest records the runtime adapter identity alongside the configured names. Running the fake path against a LIBERO YAML produces "runtime": "mindact.evaluation.fake.FakeEnvironment", so a smoke run cannot be misread as a real benchmark result.

Consistent result records

  • TrainResult, EpisodeResult, EpisodeRecord and EvaluationResult now share one freeze/validate implementation (mindact.utils.records): all frozen, keyword-only, deeply immutable and JSON-safe, each with to_dict().
  • Non-finite floats, non-string mapping keys and non-serializable values are rejected at construction rather than at write time.

Stricter configuration

  • Unknown fields rejected; booleans no longer sneak through integer fields; learning rate must be finite and positive; nested options are recursively frozen and snapshotted from the caller's dict.

CLI

  • mindact eval <yaml> --runner fake runs the full lifecycle with built-in test doubles, with --run-id, --output-dir, --episodes, --max-steps, --seed and --checkpoint overrides.
  • Adds python -m mindact.

Artifact layout

outputs/<run-id>/
├── config.yaml
├── manifest.json
├── checkpoints/
├── logs/
└── evaluation/
    ├── results.json
    └── episodes.jsonl

Verification

  • pytest -q tests/unit tests/integration — 67 passed, 2 skipped (skips are lerobot/libero not installed)
  • ruff check . — clean
  • python -m build — sdist and wheel build, both include the new modules
  • CLI smoke: --version, config-check, and eval --runner fake all produce the expected artifacts

Scope / non-goals

The fake policy and environment are contract test doubles. They are deliberately not exported from the top-level API and do not represent a trained ACT policy or a LIBERO benchmark score. Real dataset/policy/simulator adapters and the training loop follow in later PRs.

Harden configuration and provenance records, add dependency-free fake evaluation, and expose CLI smoke testing.

- Add EvaluationRunner: deterministic per-episode seeding (seed + index),
  max-step truncation, guaranteed environment close, and write-once artifacts.
- Share one freeze/validate implementation across TrainResult, EpisodeResult,
  EpisodeRecord and EvaluationResult via mindact.utils.records, so every result
  object is frozen, keyword-only and JSON-safe.
- Record runtime adapter identity in the manifest, so a fake run cannot be
  mistaken for a real LeRobot/LIBERO benchmark result.
- Add 'mindact eval --runner fake' plus python -m mindact entrypoint.
- Update README and docs to match the shipped CLI and artifact layout.
@lvyufeng
lvyufeng force-pushed the feature/mindact-pr1-evaluation branch from bdc29c1 to e3791c7 Compare August 31, 2026 15:49
@lvyufeng
lvyufeng merged commit 5db740a into master Aug 31, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant