acceptodds
Under review as a conference paper at ICLR 2027

LNPO-ATC: LENGTH-NORMALIZED POLICY OPTIMIZATION WITH ADAPTIVE TOKEN CREDIT FOR LONG-CONTEXT REASONING

Abstract

Reinforcement learning has improved reasoning in large language models, but long reasoning trajectories remain challenging under finite generation budgets. Uniformly applying a terminal length penalty can turn positive content advantages into negative update coefficients throughout a response, coupling budget control with penalties on earlier reasoning tokens. We propose Length-Normalized Policy Optimization with Adaptive Token Credit (LNPO-ATC) to address this conflict. The method combines length-normalized sequence weighting, a smooth boundary penalty, and token-level allocation based on position relative to the available budget, while retaining the content advantage at every token. Our analysis separates penalty-strength and positional contributions and derives sufficient conditions for local improvement through alignment with the true content gradient. The evaluation covers six reasoning benchmarks, long-input generalization, and allocation controls. LNPO-ATC improves the six-task macro-average score on Qwen3.6-35B-A3B by an absolute 5.53% over GSPO; ATC adds an absolute 2.27% over uniform LNPO. The ordered-versus-shuffled comparison, which also varies EOS weighting, yields an absolute 1.79% gain. At equal counts of 2,000 updates, training wall time is 1.4–1.5% higher than GSPO on the two Qwen3.6 backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.