acceptodds
Under review as a conference paper at ICLR 2027

ECA: REWARD-FREE ONLINE EFFERENCE-COPY ADAPTATION

Abstract

Learned policies are widely used to control autonomous systems, but actuator ordering, direction, or gain errors can exceed the tolerance of their feedback. Identifying and inverting the actuator map restores performance; applying an estimated inverse when no such map is present can instead degrade working feedback. We propose Efference-Copy Adaptation (ECA), which augments an identification-and-inversion core with static-map and velocity-coupled hypotheses, a rejection gate, a unit amplification bound, and an identifier hold. ECA acts under the best-supported hypothesis, withdraws when the identity or a delayed identity explains recent transitions better, decides only on informative data, and never amplifies the command it corrects. On two MuJoCo manipulators, it turns the seven-actuator permutation and the orthogonal mix, on which identification alone loses and return, into gains of and . It removes the delay losses of and return that identification causes, and converts a velocity-coupled disturbance that a map-only gate mis-fits by return into a gain of . No conservative identification control in a full grid attains ECA's recovery with a smaller worst mean loss over the four recorded preservation conditions. On a cohort trained after the method was frozen and evaluated once, recovery is met on both plants; a return loss under gain loss leaves preservation unmet on one. On a contact-rich third plant, recovery transfers but preservation does not, and a healthy-body measurement predicts the difference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.