acceptodds
Under review as a conference paper at ICLR 2027

Unifying Language Model Pre-Training and Reinforcement Post-Training with Learned Token Credit

Abstract

Pre-training and reinforcement post-training are usually treated as different algorithms. We show that both optimize one weighted autoregressive objective, differing only in the trajectory distribution and per-token learning value. This view motivates \method, a lightweight causal credit decoder inside the policy. From detached final hidden states, it learns reward, value, advantage, and importance targets; rule functions and held-out semantic chunks provide explicit supervision. No policy-sized reward model or critic is required. On matched FineWeb-Edu runs, improves the six-task zero-shot mean from 42.80 to 43.58 at 455M parameters and from 52.54 to 53.23 at 1.5B. We further reproduce a Qwen2.5-Math-7B DAPO baseline reaching 36.0% AIME-2024 avg@32 and study a preliminary trajectory reaching 37.8%. Diagnostics show why calibration matters: importance validity reaches roughly 91%, while advantage and verbal self-score correlations remain modest. The framework therefore unifies generation and judgment in one policy, while admitting learned credit only after its predictive quality is demonstrated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.