acceptodds
Under review as a conference paper at ICLR 2027

Active Value Function Learning for Long-Horizon Language Model Agents

Abstract

Long-horizon language-model agents must compact their histories to continue beyond the context window, while sparse outcome rewards make it difficult to assign credit to earlier actions. Learned critics provide stepwise value estimates, but these estimates can be inaccurate, while estimating intermediate values with additional continuations from points within a trajectory incurs substantial costs. To mitigate these issues, we introduce FOCUS, an RL method that selectively branches from self-compacted contexts to refine the critic network's value estimates. During RL training each compacted context is treated as a branchable state. The rollout controller facilitates active learning for the critic by sampling continuations from compacted states that will most reduce the critic's estimation error. These refined value estimates then score the segment between consecutive states with a TD(λ) advantage. On legal reasoning (DeonticBench), deep search (BrowseComp-Plus), and code localization (LocBench), FOCUS consistently outperforms GRPO and VinePPO with matched training steps and sampling budget, improving held-out accuracy by 7.2 points over the strongest baseline on BrowseComp-Plus (55.4 vs. 48.2) and by 4.1 points on DeonticBench (56.5 vs. 52.4).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.