Silent Compliance: Separating Constraint Echoing from Instruction Following
Abstract
Modern language models increasingly follow negative instructions well. Yet even a compliant response may restate the rule, announce compliance, or reintroduce the excluded object. Such unnecessary control language can distract users, but prior work largely evaluates only whether constraints are followed. We call this behavior constraint echoing and introduce a framework that separates visible echoing from substantive instruction following. Then we test whether echoing serves as a functional scaffold. Controlled prefix interventions show that retaining an opening echo does not help organize the answer and lowers task completion and constraint satisfaction. We next explore whether echoing is inherited from the base model or shifts at particular post-training transitions. Across sampled release transitions, SFT improves instruction following and is associated with lower echoing, whereas DPO releases exhibit consistently higher echoing. Paired likelihood tests further suggest that models can become more sensitive to a constraint without becoming more likely to echo it. Finally, we introduce Matched SFT, which trains on silent, task-successful targets selected within matched prompts and mitigates constraint echoing at low cost. Together, these results identify an overlooked failure mode, explain its functional and training-stage properties, and provide a practical way to remove it without weakening constraint following.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.