PAVIF: Reinforcement Learning for Instruction Following with the Policy as Its Own Verifier
Abstract
Instruction following is a core capability of large language models (LLMs), yet reinforcement learning for this task often relies on external verifiers to provide reward signals, introducing additional computational cost. This motivates us to investigate whether the policy can serve as its own verifier for instruction-following optimization. Our preliminary study reveals an intriguing phenomenon: verification training for the policy can also enhance its instruction-following potential. Building on this observation, we propose PAVIF, a two-stage reinforcement learning framework that internalizes verification within the policy. In the first stage, the policy is trained to verify its own responses using a self-verification dataset that we construct with responses spanning varying levels of instruction compliance. In the second stage, the policy uses its verification signals to reinforce instruction-compliant responses, turning the increased generation potential into more reliable instruction-following capability. Evaluations across six benchmarks show that PAVIF outperforms open-source baselines, with an average absolute improvement of 2.01% over the corresponding base models. Further analyses show that verification training improves the policy’s verification accuracy, while PAVIF largely preserves general capabilities and reduces per-step training time by 23.2% compared with using an external verifier of the same scale. Our code is available at https://anonymous.4open.science/r/PAVIF-310B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.