acceptodds
Under review as a conference paper at ICLR 2027

Particle Belief Representations for Reinforcement Learning under Partial Observability

Abstract

Particle filters flexibly represent complex, multimodal beliefs but expose them as weighted, unordered sets incompatible with conventional fixed-dimensional policies. We study particle-belief encodings as inputs to reinforcement-learning policies while holding belief estimation and policy optimization fixed. We compare statistical summaries, cumulant-generating-function (CGF) features, and permutation-invariant set encoders. Trainable encoders are learned end-to-end or pretrained through belief reconstruction or task-specific prediction. Across eight tasks, pooling-based set encoders and gradient CGF features are reliable end-to-end, while the Set Transformer achieves the highest performance with appropriate pretraining. Task-specific pretraining helps when its target captures decision-relevant belief structure, whereas reconstruction and latent alignment provide more selective benefits. These findings provide practical guidance for selecting and training encoders for particle-filter beliefs in continuous-space partially observable Markov decision processes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.