acceptodds
Under review as a conference paper at ICLR 2027

Generalizable Action Editor: Sharing Corrections across Tasks for Online Post-Training of Pretrained Robot Policies

Abstract

Different manipulation tasks can share local correction structure even when their goals differ. We introduce the Generalizable Action Editor (GAE), a shared policy that learns through online reinforcement learning to edit the action proposals of a frozen pretrained robot policy. Conditioning on local observations and the base proposal allows one editor to express corrections across tasks. To allocate limited training resources, we score each executed correction by its critic-estimated value gain over the same-state base proposal. This score guides CISA replay sampling and score-aware retention within ReSOP, our asynchronous multi-task learning system. A correction-support analysis characterizes how source-context coverage and compatibility of local edits shape transfer. On RoboTwin 2.0 Clean, we train one editor on 46 source tasks and evaluate it across all 50 tasks, including four held-out tasks without editor-parameter updates. Shared editing raises macro success from 90.00% to 90.84%. Four-task comparisons further favor intervention-value-based sampling over uniform and TD-error prioritization under matched learner-update budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.