acceptodds
Under review as a conference paper at ICLR 2027

Attention on Demand: Surprise-Driven Selective Attention for Reinforcement Learning

Abstract

Attention mechanisms are widely used in reinforcement learning (RL) to capture long-range temporal dependencies, yet many architectures compute temporal attention at every timestep regardless of local predictability. In partially observable environments, recurrent representations may suffice during predictable transitions, whereas longer-range temporal context can become more useful when local dynamics fail to explain incoming observations. We introduce Attention-on-Demand (AoD), a surprise-conditioned framework that adaptively weights recurrent and attention representations. AoD combines a recurrent core with causal multi-head attention and uses an auxiliary forward model to predict the next latent observation. The resulting latent prediction error serves as a local surprise signal that conditions a differentiable gate between the two representations. Across the evaluated partially observable benchmarks, AoD achieves episodic returns comparable to full-attention baselines. In the primary MemoryS7 evaluation, its mean gate activation is , corresponding to lower representational attention reliance relative to a full-attention reference. Gate activation increases from to across surprise quintiles and also rises under observation corruption. These results show that a surprise-conditioned gate can vary the attention representation's contribution with local prediction error.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.