An End-to-end Offline RL Approach to Multi-turn Agentic Alignment
Abstract
Training multi-turn agents from coarse trajectory-level feedback poses a fundamental credit-assignment problem, particularly when their interactions with users, tools, and environments are stochastic. In this paper, we propose Stochastic Trajectory Bellman Residual Minimization (Stochastic-TBRM), an end-to-end trajectory-based offline reinforcement learning approach to this problem: given static logs of agent trajectories and a single episodic alignment signal, learn a policy over the agent's sequential decision process. Under realizability, regularity, and coverage conditions, we establish a finite-sample policy guarantee for approximate minimization of the reduced trajectory objective. Controlled stochastic MDPs demonstrate value recovery and useful control from trajectory returns. Interactive-agent benchmarks evaluate deployment pipelines, whose warm-start and action-grounding components are distinguished from the effects of the learning objective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.