Do LLMs Know Where Constraints Bite? Towards Fine-Grained Reward Signals for Instruction Following
Abstract
Reinforcement learning with verifiable rewards has become the dominant post-training paradigm for instruction following. Current pipelines compress all constraint verdicts into a scalar, and broadcast the advantage uniformly across every token, while such coarse-grained credit assignment fails to localize responsible tokens for each constraint. In this paper, we investigate whether the policy's own generation distribution can be utilized for finer-grained credit. Through a pilot experiment, we find that for constraints within the model's capability, removing a constraint from the prompt causes the probabilities of exactly the tokens that implement it to drop sharply. However, for constraints beyond capability, no sampled response could satisfy them and the same contrast becomes unreliable. Based on this dichotomy, we propose LoCo, which performs constraint-level credit assignment: for constraints within the model's capability, we localize the responsible tokens by the leave-one-out contrast of each token; for constraints beyond it, we additionally provide the policy with the verbalized failure reason, so that the probability difference could localize the offending tokens. Moreover, when no sampled response satisfies all constraints, demonstration-conditioned self-distillation provides the policy with a golden response minimally edited from its best attempt to expand the model's capability boundary. Comprehensive experiments across eight instruction-following tasks confirm the effectiveness of our proposed LoCo.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.