Cross-Realization Consistency: Reducing Run-to-Run Variation in Further Training of Robot Manipulation Policies
Abstract
Even with the same data, objective, and training recipe, repeated further training of a robot policy can produce different functional changes because of training randomness. Stabilization approaches based on keeping the updated model close to the original policy can limit how much each run changes, but do not directly encourage similar functional changes across runs. We introduce Cross-Realization Consistency (CRC), which aligns the functional changes of two branches further trained from the same original policy. For independent runs, the shared change cancels in their functional difference, allowing CRC to target run-specific variation without requiring either branch to remain close to the original policy. Across eight MimicGen manipulation tasks using ACT-lite, CRC reduces functional difference by 21% on average and improves directional alignment on all tasks while maintaining average task success. Runs trained with CRC also behave more similarly during execution, in both their actions and end-effector trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.