Constraint-Decoupled On-Policy Distillation for Multi-Constraint Instruction Following
Abstract
Multi-constraint instruction following requires a model to satisfy every constraint in an instruction simultaneously. We show that the dominant failure mode of current LLMs is partial satisfaction: the all-satisfy rate (ASR) collapses as constraints accumulate, while the per-constraint satisfaction rate stays above 89%. Neither scalar RLVR nor standard on-policy distillation repairs this, because both discard constraint-level structure: RLVR compresses all constraint evaluations into a single reward, and on-policy distillation conditions the teacher on the full instruction, so the student inherits the teacher's own compositional decay. We propose Constraint-Decoupled On-Policy Distillation (CD-OPD), which decouples the instruction into single-constraint prompts, scores each constraint's token-level effect against an unconstrained base, and combines them through a differentiable min-pooling operator. This operator acts as a soft veto—a token violating any constraint is suppressed—while a full-instruction anchor preserves fluency, and the resulting target is distilled into the student along its own on-policy trajectories. Across three teacher–student scales(0.6/1.7/4B) and four benchmarks, CD-OPD improves average ASR over the base student by 4.48%/2.54%/3.57%, reduces the FollowBench decay from 36.6% to 24.5% over one to five constraints, while training substantially faster than RLVR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.