Unifying Language Model Pre-Training and Reinforcement Post-Training with Learned Token Credit
Abstract
Pre-training and reinforcement post-training are usually treated as different algorithms. We show that both optimize one weighted autoregressive objective, differing only in the trajectory distribution and per-token learning value. This view motivates \method, a lightweight causal credit decoder inside the policy. From detached final hidden states, it learns reward, value, advantage, and importance targets; rule functions and held-out semantic chunks provide explicit supervision. No policy-sized reward model or critic is required. On matched FineWeb-Edu runs, improves the six-task zero-shot mean from 42.80 to 43.58 at 455M parameters and from 52.54 to 53.23 at 1.5B. We further reproduce a Qwen2.5-Math-7B DAPO baseline reaching 36.0% AIME-2024 avg@32 and study a preliminary trajectory reaching 37.8%. Diagnostics show why calibration matters: importance validity reaches roughly 91%, while advantage and verbal self-score correlations remain modest. The framework therefore unifies generation and judgment in one policy, while admitting learned credit only after its predictive quality is demonstrated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.