When Atomic Constraint Application Does Not Compose: Diagnosing Constraint Composition Failure in LLMs
Abstract
Recall does not establish that a model can apply a constraint. We define application-conditioned residual Constraint Composition Failure (CCF): failure to construct a feasible joint assignment after exact rule representation and isolated application have been demonstrated. Our gate exhaustively tests active operator inputs using the joint task's canonical context and complete-assignment interface. In 50-seed Matched Constraint Ring confirmations, residual failure at L2/L4 is .740/.935 for Qwen3.7-Max and .762/.867 for SenseNova-6.7-Flash-Lite. On seeds admitted at both lengths, Qwen increases by 21.7 points (paired 95% CI [6.5,37.0], p=.0129), whereas SenseNova's 7.9-point increase is inconclusive ([-7.9,23.7], p=.5078). A matched final-only repair instruction yields no reliable reduction. Repository-grounded evidence is model-family specific: DeepSeek-V4-Flash fails on 17/32 Updo and 4/24 Pebble control-success strict runs, whereas a jointly frozen Qwen extension succeeds on all 50 Updo runs but fails all 37 eligible Pebble runs. Deadlock and CodePlan audits delimit adjacent application, evaluation, and ordinary reliability failures. CCF is therefore a controlled diagnosis of an atomic-to-joint capability gap, not a universal depth law, model-wide deficit, software-task prevalence estimate, or renamed combinatorial explosion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.