acceptodds
Under review as a conference paper at ICLR 2027

Agentic Policy Improvement through Value-guided Informative Search

Abstract

Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models. However, its learning signal is limited by the trajectories discovered by the current policy. When group rollouts repeatedly follow similar or low-value paths, the resulting verified signal provides little guidance for learning better behavior. Existing search-based rollout methods can increase trajectory diversity, but diversity alone neither guarantees a better rollout policy nor ensures that search improvements are transferred back to the model. We therefore study how to make search during RLVR a principled policy-improvement operator. We propose APIVIS, a rollout-and-learning framework that applies value-guided search to domain-native language-model decisions and combines searched and direct trajectories within each rollout group. APIVIS further distills search-improved decisions, preserving an improvement-aligned learning signal even when group-relative returns are uninformative. We prove statewise and trajectory-level policy improvement under exact values, together with an explicit degradation bound under bounded value errors. Across established benchmarks in mathematical reasoning and multi-turn interaction, APIVIS outperforms established RLVR and search-based baselines, demonstrating the effectiveness of search as a learnable policy-improvement operator.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.