acceptodds
Under review as a conference paper at ICLR 2027

Constraint-Decoupled On-Policy Distillation for Multi-Constraint Instruction Following

Abstract

Multi-constraint instruction following requires a model to satisfy every constraint in an instruction simultaneously. We show that the dominant failure mode of current LLMs is partial satisfaction: the all-satisfy rate (ASR) collapses as constraints accumulate, while the per-constraint satisfaction rate stays above 89%. Neither scalar RLVR nor standard on-policy distillation repairs this, because both discard constraint-level structure: RLVR compresses all constraint evaluations into a single reward, and on-policy distillation conditions the teacher on the full instruction, so the student inherits the teacher's own compositional decay. We propose Constraint-Decoupled On-Policy Distillation (CD-OPD), which decouples the instruction into single-constraint prompts, scores each constraint's token-level effect against an unconstrained base, and combines them through a differentiable min-pooling operator. This operator acts as a soft veto—a token violating any constraint is suppressed—while a full-instruction anchor preserves fluency, and the resulting target is distilled into the student along its own on-policy trajectories. Across three teacher–student scales(0.6/1.7/4B) and four benchmarks, CD-OPD improves average ASR over the base student by 4.48%/2.54%/3.57%, reduces the FollowBench decay from 36.6% to 24.5% over one to five constraints, while training substantially faster than RLVR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.