acceptodds
Under review as a conference paper at ICLR 2027

Not Every Observation Should Update Belief Equally

Abstract

Partially observable reinforcement learning requires agents to maintain an internal state from incomplete observations. Although recurrent agents can attenuate incoming information through learned gates, revision strength is typically learned only implicitly through end-to-end control. Yet observation surprise and decision relevance are distinct: a surprising observation may be irrelevant, while a plausible cue may alter the preferred action. We introduce BeliefGate, which makes assimilation strength an explicit decision variable supervised by the downstream control utility of alternative belief revisions. During training, paired interventions compare candidate revision strengths from the same simulator state under matched future randomness, and the intervention selected by penalized utility supervises a continuous deployment-time gate. We evaluate diagnostic POMDPs with matched-surprise and same-token context tests, together with long-memory and observation-shift benchmarks. BeliefGate achieves the strongest overall and unseen-shift aggregates, while specialized memory baselines remain stronger on some memory-centric tasks. It also provides the strongest non-oracle relevance separation and AUROC under matched surprise, with the smallest relative degradation across observation shifts. Representation probes further show stronger decision-relevant content and lower distractor retention. These results frame observation assimilation as a decision-relevant control problem rather than a generic filtering operation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.