LLMs Can Follow Instructions, But Not Many at Once: Multiplicative Collapse in Compositional Instruction Following
Abstract
Large language models follow a single explicit instruction reliably, but deployed prompts rarely contain just one. We ask what happens as instructions accumulate: how fast the chance of satisfying all of them declines, what sets the rate, and whether effort at inference time can slow it. We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that varies the number of simultaneous constraints from 1 to 12 and checks every constraint with a deterministic verifier, with no LLM judge at any stage. Across 19 models, 36 constraint types, and 472,019 checks, per-constraint pass rates decline gently while joint success collapses: at , models pass 45% of individual constraints but satisfy all eight on only 9.6% of prompts. No pair of constraints interferes: failures of different constraints are nearly uncorrelated, so the risk of failure compounds across constraints, and no choice or arrangement of constraints avoids the collapse. The per-constraint rates themselves follow a clear hierarchy: constraints that must be tracked across the whole output lose their single-constraint accuracy about twice as fast as lexical ones, and constraints settled by a single decision degrade least. When constraints cannot all be met, models keep the ones that hold up best under load. Planning before generation does not delay the collapse; self-correction and best-of-5 sampling delay it by one or two constraints. Only four of the 19 models satisfy every constraint on a majority of prompts at , and 12 lose that majority at three constraints or fewer. We will release all verifiers, probes, model outputs, and code for reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.