acceptodds
Under review as a conference paper at ICLR 2027

Contextual Credit Diffusion for LLM Post-training

Abstract

LLM post-training has long relied on a single objective at each position: optimizing a weighted log-likelihood of a single reference token in the vocabulary. This reference token is either the sampled action in RL methods such as PPO and GRPO, or the ground-truth label in SFT. The loss can be written in the unified form (L = -E_t [ w_t cdot log pi_ theta(a_t^ref mid s_t) ] ), where (w_t ) is an advantage in RL methods or a constant in SFT. However, solving a reasoning problem often admits multiple valid paths that all lead to the correct answer, yet the forced-reference objective locks the model onto a single path at each step and penalizes alternative tokens in the vocabulary that could also lead to a valid answer. We propose **Contextual Credit Diffusion (CCD)**. CCD diffuses the reference token’s credit across the vocabulary according to their contextual interchangeability, measured by the logit difference in the model’s next-token distribution under the current context, replacing the single-reference constraint with a dense reward. This offers a potential solution to both the credit assignment problem in RL and the one-hot label problem in SFT. CCD consistently outperforms RL and SFT baselines on six reasoning benchmarks, achieving significant gains. Furthermore, a token-replacement experiment shows that replacing the reference token with candidate tokens whose diffused credit is high yields the same final answer in most cases, confirming that the reward captures contextual interchangeability and identifies alternative valid paths.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.