When One More Constraint Is Added: Decomposing Multi-Constraint Failure in Instruction-Following Evaluation
Abstract
Multi-constraint instruction-following benchmarks usually reduce a response to one score: did it satisfy every constraint at once? When that score falls as constraints are added, it does not say why. Did the model fail the new constraint, lose one it had already satisfied, or did the evaluation change its judgment while the behaviour stayed the same? We follow each constraint as the prompt grows, separating two events: satisfying the newly added constraint, and breaking one that previously passed. On FollowBench this needs two automated steps, and both are fragile. Software extracts the text of each added constraint, and an LLM judge decides whether a response satisfies it. Two plausible judging protocols disagree by as much as 55.9 percentage points on identical outputs. On a slice of Llama responses with rule-based checkers, the judge marks 30.2% of rule-passing constraints as failed, and 13 of 14 constraints that failed and then passed again were judge errors, not real recoveries. Human annotators confirm 297 of the 392 tracked constraints, counting every disagreement or ambiguous case as not confirmed. Under our frozen analysis, constraints already satisfied show no detectable steady rise in breakage across four instruction-tuned models, though rises up to 4.7 points per added constraint remain compatible with the data, and dropping the 95 unconfirmed constraints gives the same answer. We then repeat the question on ManyIFEval, where instructions are listed separately and checked by a program, so neither step is needed. Following one instruction while four more are added, its satisfaction changes by between -1.44 and +1.20 points per added instruction, with every 95% interval inside [-2.36, +2.64]; only Qwen declines detectably. Yet the chance that an instruction satisfied at one step fails at the next rises in two models. How often a constraint is satisfied and how often a satisfied constraint is lost give different answers about the same responses, even under exact scoring. Multi-constraint evaluation therefore depends not only on how many constraints are satisfied, but on which change is measured and how it is judged.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.