Equivalent Rewards, Unequal Interfaces: Reward Coordinates in History-Conditioned RL
Abstract
History-conditioned reinforcement learning agents use rewards both to define an objective and as inputs for online adaptation. Does preserving the objective also preserve adaptation? We distinguish objective, information, and policy-interface equivalence, and study their separation using interventions that alter only the rewards visible to the actor. Across five independently trained AMAGO policies in each of two continuous-control task families, positive-affine reward transformations alter actions and original-task returns. HalfCheetah exhibits a non-monotone, training-coordinate-specific response; Ant-Fwd-Back reproduces closed-loop sensitivity and a direct, reversible reward-to-action pathway. Teacher-forced replay localizes effects before environmental feedback, while HalfCheetah activation patching identifies structured, action-relevant changes not captured by the decoded task axis. Recurrent RL agents also exhibit strong coordinate sensitivity, whereas Decision Transformer controls show weaker sensitivity consistent with their weaker reliance on trajectory-specific reward evidence. Zero and donor interventions further connect coordinate sensitivity to reward reliance. These results demonstrate reward-interface dependence in learned adaptation: equivalent objectives and information need not produce equivalent behavior when reward is part of the policy interface.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.