acceptodds
Under review as a conference paper at ICLR 2027

ReLPO: Resolution-Lifted Policy Optimization for RL Post-Training of Autoregressive–Diffusion Language Models

Abstract

Autoregressive (AR)–diffusion hybrid language models combine revisable latent refinement with irreversible token realization, creating a distinct challenge for reinforcement learning (RL) post-training in assigning outcome credit across both forms of decision making. We observe that hybrid generation naturally induces a native relaxation hierarchy. Each partially resolved state defines a completion fiber containing all terminal sequences consistent with its committed tokens, while its unresolved latent state determines a continuation distribution over this fiber. Latent refinement reshapes this distribution while preserving possible completions, whereas autoregressive realization progressively contracts the fiber toward a single discrete output. Building on this structure, we introduce ReLPO (Resolution-Lifted Policy Optimization), which propagates terminal outcome information through the hierarchy using a Resolution-Lifted Outcome Value (RLOV). RLOV defines a soft potential over partially resolved states and is learned from completed rollouts and terminal verifier rewards through a consistency relation spanning both refinement and resolution transitions. The resulting shared potential guides continuous latent refinement and autoregressive realization through a coupled policy improvement operator, allowing terminal outcomes to coordinate refinement, deferral, and token commitment throughout generation. Across eight mathematical reasoning and code-generation benchmarks, ReLPO improves the shared Evo backbone from 52.5 to 59.7 average accuracy, compared with 56.4 for the strongest of eight RL baselines. Using separate potentials for the two branches reaches 57.7 under matched capacity, demonstrating the benefit of coupling refinement and realization through a shared potential. Under a unified protocol across six diffusion and AR–diffusion backbones, ReLPO improves every model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.