acceptodds
Under review as a conference paper at ICLR 2027

RegPO: Leveraging Regret-Based Feedback for Agentic Post-Training with Tree-Structured Sampling

Abstract

We introduce RegPO, a policy optimization method that uses regret-based feedback for step-level credit assignment in post-training large language model (LLM) agents. Specifically, motivated by the decomposition of regret for learning in Markov decision processes (MDPs) into action suboptimality gaps, RegPO develops sampled-search proxies for regret-based advantages at selected prefixes, leveraging tree-based sampling. These per-step regret proxies are constructed from sampled rollout subtrees by comparing each action’s estimated value with the highest estimate at the same prefix. We show that, in finite deterministic tree-structured MDPs, existing Tree-GRPO’s return-based advanatages can decrease the probability of selecting an optimal action during learning. This occurs when the action’s value under the current policy is lower than that of a suboptimal action. Experiments on these MDPs then show how the different advantage constructions affect the learning dynamics of policy optimization, e.g., which actions remain likely to be sampled during training. An ALFWorld analysis further examines how action value rankings relate to action selection, task-relevant intermediate-state reachability, and continuation execution in agentic settings. We also evaluate RegPO extensively on four agentic benchmarks spanning robotic manipulation, web interaction, browser control, and embodied reasoning. Under the benchmark-specific reporting protocol, the evaluated RegPO configurations attain the highest scores among the compared methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.