acceptodds
Under review as a conference paper at ICLR 2027

One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness

Abstract

Instruction-tuned large language models produce thorough responses when prompted normally. We show that this helpfulness is fragile. Adding a trivial constraint, such as a ban on a single punctuation mark or on the word “the”, costs **17-48%** of comprehensiveness on average in pairwise evaluation of seven models, and all three judges detect the loss. The loss is in content rather than presentation: content criteria degrade - more than surface criteria, and a blinded human evaluation reproduces the pattern. The trigger is the constraint's presence in the prompt, not the effort of obeying it. Four converging tests support this, including bans on words that essentially never occur, which still cost 10-21%, and decoder-level enforcement of the identical constraint with a clean prompt, which removes almost all of the loss (26.0% 4.5%). The response-strategy decision is made before generation, where a linear probe can detect it (-), and it is reversible at a single mid-network layer: steering along a difference-of-means direction with a compliance-preserving decoder mask recovers 64-82% on Qwen-2.5-7B at 100% constraint satisfaction, with no training, and at most 14% at the six other depths tested. Beyond the best steering dose, responses stay near full length while comprehensiveness recovery falls by 18 points, so the judge is not rewarding length. The collapse extends to deployment constraints (14-31%), independent scoring underestimates it by -, and LoRA fine-tuning on 80 examples recovers 65-86%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.