acceptodds
Under review as a conference paper at ICLR 2027

Self-Confirming Superposition Traps in Reinforcement Learning

Abstract

Reinforcement learning (RL) trains representations on data selected by the agent's policy and updates that policy using the returns they support. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal for the policy's own data. Specifically, we identify a self-confirming superposition trap, in which every optimal code shares features that rarely co-activate under the current policy but interfere after an alternative action. The resulting control error lowers that action's return and reinforces its avoidance, although choosing it would earn more after representation adaptation at the same capacity. In a tied two-step model, we characterize which feature counts and representation dimensions admit a trap. We also show separately that continued visitation, equal feature frequencies, and independent controller learning need not prevent it. Because rarely chosen actions contribute little to the fitting objective, optimal fitting can tolerate errors large enough to reverse their return advantage. We derive a bound on this distortion in a broader finite-action model and use it to prove a sufficient replay condition: retaining enough training weight on the best separately adapted action preserves its return advantage despite residual representation error. Experiments with neural PPO show that agents initialized toward different actions develop different interference patterns and opposite mean return rankings at the same capacity, with corresponding differences in their final policies. Motivated by the replay condition, we test whether preserving access to neglected training states can improve control. Providing such access reduces measured interference and improves sequential return, with return gains even when the encoder is frozen. We extend these tests to MiniGrid and DMControl, where training-state access, replay reweighting, and overlap penalties can improve control. In DreamerV3–Crafter, protecting world-model fitting weight also improves cumulative training scores at unchanged capacity. Our code is available at https://anonymous.4open.science/r/self-confirming-superposition-traps-SCST.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.