Particle Belief Representations for Reinforcement Learning under Partial Observability
Abstract
Particle filters flexibly represent complex, multimodal beliefs but expose them as weighted, unordered sets incompatible with conventional fixed-dimensional policies. We study particle-belief encodings as inputs to reinforcement-learning policies while holding belief estimation and policy optimization fixed. We compare statistical summaries, cumulant-generating-function (CGF) features, and permutation-invariant set encoders. Trainable encoders are learned end-to-end or pretrained through belief reconstruction or task-specific prediction. Across eight tasks, pooling-based set encoders and gradient CGF features are reliable end-to-end, while the Set Transformer achieves the highest performance with appropriate pretraining. Task-specific pretraining helps when its target captures decision-relevant belief structure, whereas reconstruction and latent alignment provide more selective benefits. These findings provide practical guidance for selecting and training encoders for particle-filter beliefs in continuous-space partially observable Markov decision processes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.