acceptodds
Under review as a conference paper at ICLR 2027

Belief-Centric Reinforcement Learning for LLM Agents in POMDP Tasks

Abstract

Decision making under uncertainty is a central challenge for large language model (LLM) agents (e.g., robotics). Partially Observable Markov Decision Processes (POMDPs) provide a principled framework for modeling decision-making from incomplete observations. In POMDP tasks, agent’s performance depends on two mutually coupled capabilities: how to estimate the state-relevant belief information, and how to disentangle the estimation of failure beliefs and assign the accurate credit to actions. To address this challenge, this paper proposes BeliefRL, a belief-centric RL framework that jointly adapts belief-state estimation and policy learning. BeliefRL first maintains a structured belief state over established, ruled-out, and unresolved information, continually adapts the belief-state model using simulator-grounded supervision. These state-reference discrepancies are further measured as a reliability signal to modulate policy updates. BeliefRL is validated on four partially observable tasks in VirtualHome and Overcooked. Compared with POAD, the strongest overall baseline, BeliefRL improves averaged return by 44.1% and normalized training AUC by 34.8%. Ablations support the complementary benefits of training-time belief adaptation and reliability-aware policy optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.