Adapting Learned Robot Simulators: Policy Visitation Beats Error Magnitude
Abstract
Policies trained inside learned robot simulators exploit the simulators' errors: on the public NeRD model of the Ant robot, jump policies score higher in the model than in the physics engine it imitates. The prevailing remedy is to adapt the simulator with data collected where the model errs most, but in learned robot simulators this has not been tested against collecting where the policy goes. We provide that test with two tools: a controlled testbed in which seven acquisition rules, including an adversary given the model's true error, adapt the released NeRD models under identical data budgets and fine-tuning; and a paired-perturbation diagnostic that checks whether a model reproduces the engine's contact-event changes. The adversary finds errors – larger than random actions, yet yields policies no better than random data, whereas collecting along the trajectories of model-trained policies recovers – of the return of a policy trained in the engine on the jump task. Fine-tuning that cuts one-step error fivefold leaves most contact-event changes unreproduced, and the rule with the lowest boundary error trains worse policies than policy-trajectory collection. The value of adaptation data thus depends on whether the policy will visit its states, not on how wrong the model is there, and prediction accuracy, contact fidelity, and policy performance must be evaluated separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.