Anchored Anytime Value Learning for Long-Horizon Agent Reinforcement Learning
Abstract
Long-horizon agent training calls for anytime value learning that decouples from trajectory-completion constraints. We introduce a barrier-synchronized streaming framework that optimizes in-flight trajectory chunks, eliminating straggler-induced idle time while preserving chunk-level policy freshness. Value learning over nonterminal chunks, however, corrupts along two axes: it inflates the magnitude variance of the bootstrap propagated within a trajectory, and it scrambles the ordering of values across concurrent trajectories. Mechanistic analysis shows: by unfolding the anisotropic observation manifold into a higher-resolution belief space, thinking widens task-level (between-task) separation far faster than it changes surface-level (within-task) variation, boosting the Fisher signal-to-noise ratio of the value representation. We then propose Anchored Anytime Value Learning (A2VL): (1)longitudinal magnitude anchoring shrinks noisy boundary bootstraps toward the low relative-variance belief-state value to limit temporal error propagation; and (2)transverse ordering anchoring imposes pairwise ordering constraints across concurrent trajectories using the well-separated belief-state ordering as a detached reference. Across WebShop, ALFWorld, and Crafter, A2VL stabilizes anytime training across chunk granularities, matching episode-synchronous task performance while substantially improving wall-clock throughput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.