acceptodds
Under review as a conference paper at ICLR 2027

POMDP Interfaces in Agentic Reinforcement Learning: A Framework-Driven Review

Abstract

Agentic reinforcement learning extends language-model training to learning through long-horizon interaction. Existing surveys organize the field by capabilities, applications, or algorithm families, leaving less explicit how choices in memory, action abstraction, feedback, and experience collection interact. These dependencies matter because the same mechanism can improve learning in one configuration yet impair it in another. In this review, we develop a POMDP-based perspective that traces how task structure is translated into an implemented policy-gradient update. The resulting Four-Gap POMDP Framework organizes the literature around four interfaces: Belief, Granularity, Signal, and Rollout. Belief and Granularity determine the information states and decision units of the induced process; Signal and Rollout determine how its gradient is estimated. We synthesize research across these interfaces, examining their operating assumptions, characteristic failures, and six pairwise couplings through shared states, decision events, and distributions. A finite-difference formulation characterizes configuration-dependent interactions, while an exact Signal–Rollout identity separates credit error, correction residuals, and their joint contribution to update bias. Controlled ALFWorld studies illustrate how outcome anchoring and collection lag shape learned-evaluator failure. This perspective connects research directions through the decision objects they jointly construct, providing a basis for comparing methods and identifying compatible designs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.