acceptodds
Under review as a conference paper at ICLR 2027

ReMPO: Behavior-Consistent Latent-Depth Policy Optimization for Search Agents

Abstract

Looped language models (LoopLMs) enable multi-step latent reasoning by repeatedly applying shared computation, offering a promising way to increase reasoning depth without increasing model parameters. However, their recurrent computation also creates a mismatch with standard reinforcement learning, which is defined around the final policy output. Recent work has taken an important first step by extending outcome supervision to intermediate predictions, allowing internal computation to receive learning signals. Yet reinforcement learning (RL) is fundamentally driven by the behavior policy that generates the trajectory: assigning credit to intermediate states does not by itself make them part of the policy that produces actions. We therefore introduce Readout-Marginalized Policy Optimization (ReMPO), which constructs the behavior policy by marginalizing token predictions across recurrent readouts. ReMPO uses the same readout-marginalized policy for both trajectory generation and policy optimization, allowing final-answer rewards to jointly train token prediction and readout selection. A terminal anchor and regularization stabilize this learned policy. Using only final-answer rewards, we train Ouro-1.4B and Ouro-2.6B as search agents for question answering. ReMPO outperforms the state-of-the-art method on all seven datasets at both model sizes; at 1.4B, it improves the average exact match by 11.30 percentage points and held-out multi-hop performance by 17.76 points. It also yields more effective two-search trajectories, with second retrievals recovering answer-bearing evidence missed by the first.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.