WHETHER VS. WHEN: SEPARATING SHORTCUT FORMATION FROM EVIDENCE-TIMING SENSITIVITY IN VALUE-BASED ONLINE REINFORCEMENT LEARNING
Abstract
Deep reinforcement learning agents can overweight early experience, a phenomenon studied as primacy bias, but whether a shortcut forms and how strongly its formation depends on the timing of disambiguating evidence are distinct properties that this framing can conflate. We separate them in a controlled visual task where a salient, spurious feature and a subtler, reward-relevant feature agree except at planted evidence cells. Holding the content and quantity of evidence fixed, we relocate the same 40 evidence-cell visits from the first to the last 400 steps of a 4000-step value-based online Double DQN training window. This temporal shift changes shortcut formation from 0 of 30 to 30 of 30 paired seeds, with a mean paired margin difference of −6.17 (95% CI [−6.41, −5.93]). The split survives interventions targeting action coverage, exploration, evidence dose, integration time, replay-sampling exposure, and optimizer-update count, although inward placement and closer matching of integration time attenuate it. The effect also persists across a smaller vision transformer, a structurally distinct corridor environment, and fully organic, action-driven visitation (0/30 vs. 30/30), though realized exposure there is endogenous and unequal rather than matched. In contrast, matched fixed-dataset supervised and bootstrapped-TD learners show substantially smaller timing effects, while a matched on-policy PPO agent shows no timing-dependent split despite forming the same shortcut. Removing experience replay also closes the split, but batch size and gradient variance change simultaneously, so this does not identify replay as the unique cause. A rewind intervention further points to the role of accumulated network state in how evidence is incorporated. Together, these results show that shortcut formation is not unique to online reinforcement learning, but its dependence on evidence timing is substantially amplified in the value-based online-RL regime studied here, while leaving the precise mechanism unresolved.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.