acceptodds
Under review as a conference paper at ICLR 2027

Cross-State Action-Effect Supervision: Which Correspondence Matters?

Abstract

Cross-state action-effect supervision compares the same action substitution at two starting states. When learners receive the same branch labels, which state correspondences contribute to its training benefit? We decompose the auxiliary loss into common and differential effect residuals and use controls that preserve action identity and match auxiliary-gradient strength. In a controlled dynamical system, reassigning position and velocity within actuator modes retains improvements over paired supervision; exact alignment establishes no advantage over this control. We then vary whether mode and gate affect motion while preserving their visibility and matching permutation cycle lengths. Under the original physics, mode-preserving re-pairing reduces common-active effect error by 23.21% relative to cross-effect-class re-pairing. A paired endpoint comparison establishes dependence on the physical condition, and independent intermediate-condition cohorts show a 14.26% reduction. Cross-class re-pairing also improves over the paired baseline, so preserving mode changes the size of the benefit rather than determining whether any benefit exists. Lower differential error can coexist with greater false response or longer-horizon effect error. Native-system studies also show that response and absolute prediction gains need not coincide, and establish no general control advantage. These findings motivate action-preserving correspondence controls alongside absolute-error and readout checks for relational supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.