acceptodds
Under review as a conference paper at ICLR 2027

Information-Time Proximal Policy Optimization

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has substantially improved the reasoning capabilities of Large Language Models (LLMs). However, existing methods typically measure temporal progression by token count, despite the highly non-uniform information flow along autoregressive trajectories. We propose Information-Time Proximal Policy Optimization (InfoPPO), which measures temporal distance by accumulated state-wise information rather than raw token count. This provides a common basis for temporal credit propagation and policy updates. InfoPPO enables non-trivial discounting in long-horizon reasoning, retaining effective-horizon contraction in information time while mitigating excessive attenuation of terminal supervision. Moreover, the information-time policy-improvement analysis naturally leads to a state-dependent update constraint, which we implement through adaptive clipping. It scales the update range with local information density, enabling targeted policy updates while preserving proximal control. Theoretically, we extend performance-difference and policy-improvement analyses to the information-time Markov Decision Process (MDP), deriving a policy-improvement lower bound. We further connect this analysis to practical LLM policy optimization by relating state-wise information density to local policy movement, while also providing theoretical grounding for the adaptive update mechanism. Experiments on Qwen3 models demonstrate consistent gains over competitive baselines across five challenging competition-style mathematical reasoning benchmarks. InfoPPO also maintains stable accuracy and response length across non-trivial discount settings under which token-time PPO deteriorates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.