Ask first, code second.
LIBERO's eval script reports success rate per task suite. That is the right primary number. The gap I keep hitting is downstream: a growing number of papers report a LIBERO success rate under some perturbation and compare it to the clean number from a different run, different seeds, sometimes a different commit. The difference they publish is then partly the perturbation and partly run-to-run spread, and there is no way for a reader to separate the two.
What would close it, inside your existing structure and without changing default behaviour:
- A
--control-arm flag, off by default. When set, the script replays the same seed list with the perturbation hook disabled.
- Two extra keys in the results JSON:
success_rate_control and n_episodes.
- A Wilson interval on both rates, so the reader can see whether n supported the comparison.
Measured reason to bother: on 4 open policies and 16 perturbation families, the control arm fires at a family-dependent rate, and subtracting it removes the apparent effect in 13 of 16 families at n=50. If that holds on LIBERO tasks specifically, a lot of published gaps are inside the noise.
I would write this, keep it to the eval path, and add tests. But it is your benchmark and a flag in the official script is a convention, not a patch, so I am asking before writing rather than after.
Also happy to be told this belongs in each user's own wrapper. That is a clear answer.
Ours, for context and so you can check the claim rather than take it: https://github.com/provael/provael - simulation-only, no hardware results.
Ask first, code second.
LIBERO's eval script reports success rate per task suite. That is the right primary number. The gap I keep hitting is downstream: a growing number of papers report a LIBERO success rate under some perturbation and compare it to the clean number from a different run, different seeds, sometimes a different commit. The difference they publish is then partly the perturbation and partly run-to-run spread, and there is no way for a reader to separate the two.
What would close it, inside your existing structure and without changing default behaviour:
--control-armflag, off by default. When set, the script replays the same seed list with the perturbation hook disabled.success_rate_controlandn_episodes.Measured reason to bother: on 4 open policies and 16 perturbation families, the control arm fires at a family-dependent rate, and subtracting it removes the apparent effect in 13 of 16 families at n=50. If that holds on LIBERO tasks specifically, a lot of published gaps are inside the noise.
I would write this, keep it to the eval path, and add tests. But it is your benchmark and a flag in the official script is a convention, not a patch, so I am asking before writing rather than after.
Also happy to be told this belongs in each user's own wrapper. That is a clear answer.
Ours, for context and so you can check the claim rather than take it: https://github.com/provael/provael - simulation-only, no hardware results.