POMDP Interfaces in Agentic Reinforcement Learning: A Framework-Driven Review
Abstract
Agentic reinforcement learning extends language-model training to learning through long-horizon interaction. Existing surveys organize the field by capabilities, applications, or algorithm families, leaving less explicit how choices in memory, action abstraction, feedback, and experience collection interact. These dependencies matter because the same mechanism can improve learning in one configuration yet impair it in another. In this review, we develop a POMDP-based perspective that traces how task structure is translated into an implemented policy-gradient update. The resulting Four-Gap POMDP Framework organizes the literature around four interfaces: Belief, Granularity, Signal, and Rollout. Belief and Granularity determine the information states and decision units of the induced process; Signal and Rollout determine how its gradient is estimated. We synthesize research across these interfaces, examining their operating assumptions, characteristic failures, and six pairwise couplings through shared states, decision events, and distributions. A finite-difference formulation characterizes configuration-dependent interactions, while an exact Signal–Rollout identity separates credit error, correction residuals, and their joint contribution to update bias. Controlled ALFWorld studies illustrate how outcome anchoring and collection lag shape learned-evaluator failure. This perspective connects research directions through the decision objects they jointly construct, providing a basis for comparing methods and identifying compatible designs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.