feat: add deterministic evaluation runner - #2772
Merged
Merged
Conversation
Harden configuration and provenance records, add dependency-free fake evaluation, and expose CLI smoke testing. - Add EvaluationRunner: deterministic per-episode seeding (seed + index), max-step truncation, guaranteed environment close, and write-once artifacts. - Share one freeze/validate implementation across TrainResult, EpisodeResult, EpisodeRecord and EvaluationResult via mindact.utils.records, so every result object is frozen, keyword-only and JSON-safe. - Record runtime adapter identity in the manifest, so a fake run cannot be mistaken for a real LeRobot/LIBERO benchmark result. - Add 'mindact eval --runner fake' plus python -m mindact entrypoint. - Update README and docs to match the shipped CLI and artifact layout.
lvyufeng
force-pushed
the
feature/mindact-pr1-evaluation
branch
from
August 31, 2026 15:49
bdc29c1 to
e3791c7
Compare
7 of 39 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds the first working piece of MindAct beyond the skeleton: a deterministic, dependency-free evaluation runner with immutable experiment provenance. No LeRobot, LIBERO, ACT or training-loop code is introduced here; this PR establishes the contract those integrations will plug into.
What's in it
Evaluation runner (
src/mindact/evaluation/runner.py)config.seed + N, so a run is reproducible from its manifest alone.reset(seed=...)->predict(observation)->step(action)untildoneorevaluation.max_steps; hitting the step limit setstruncated=True.environment.close()always runs in afinallyblock, including when the policy raises.done, mappinginfo, JSON-serializable metadata.Write-once provenance
manifest.jsonand result files are created via a same-directory temp file,fsync, thenos.link()for exclusive creation.evaluation/results.json) and per-episode records (evaluation/episodes.jsonl) are separate files, so completing an evaluation never mutates the manifest."runtime": "mindact.evaluation.fake.FakeEnvironment", so a smoke run cannot be misread as a real benchmark result.Consistent result records
TrainResult,EpisodeResult,EpisodeRecordandEvaluationResultnow share one freeze/validate implementation (mindact.utils.records): all frozen, keyword-only, deeply immutable and JSON-safe, each withto_dict().Stricter configuration
optionsare recursively frozen and snapshotted from the caller's dict.CLI
mindact eval <yaml> --runner fakeruns the full lifecycle with built-in test doubles, with--run-id,--output-dir,--episodes,--max-steps,--seedand--checkpointoverrides.python -m mindact.Artifact layout
Verification
pytest -q tests/unit tests/integration— 67 passed, 2 skipped (skips arelerobot/liberonot installed)ruff check .— cleanpython -m build— sdist and wheel build, both include the new modules--version,config-check, andeval --runner fakeall produce the expected artifactsScope / non-goals
The fake policy and environment are contract test doubles. They are deliberately not exported from the top-level API and do not represent a trained ACT policy or a LIBERO benchmark score. Real dataset/policy/simulator adapters and the training loop follow in later PRs.