Diagnosing & Mitigating Constraint Induced Response Collapse in Instruction-Tuned LLMs
Abstract
Instruction-tuned language models are trained to give thorough, well-organized answers. Recent work showed that this behavior is fragile: banning a single character, such as the em dash, can remove up to 48% of an answer's substantive content. This _constraint-induced response collapse_ has been ascribed to shallow template following, but its mechanism remains unclear. We show that lexical bans are only one of many triggers: across seven models, everyday constraints on tone, formatting, and delivery cut comprehensiveness by up to 45%, even though the model obeys them. Controlled ablations identify the trigger as _form primers_, phrases in the prompt that refer to the form of the response rather than its content. Studying each post-training stage of the open OLMo family, we trace the collapse to reward-based post-training rather than template following: it emerges at preference optimization and is largest after reinforcement learning from verifiable rewards, two stages whose rewards do not measure content. A _single direction_ in the model's activations controls the collapse: adding it to unconstrained prompts induces collapse, and subtracting it from constrained prompts restores most of the lost content with the constraint still obeyed, mitigating the collapse at inference time. Because pointwise judges miss about half of this loss, we release a 160-question benchmark that scores every answer against the same model's unconstrained answer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.