Why Do Language Models Take Shortcuts? A Mechanistic Account of Heuristic–Constraint Competition
Abstract
Large language models (LLMs) can solve complex reasoning problems yet fail on decisions that humans find trivial, for example, recommending walking to a nearby car wash while overlooking that the car itself must be brought there. These failures expose a fundamental limitation of current reasoning systems: possessing the right knowledge does not guarantee using it when it matters. Salient surface heuristics can override available constraints, and simply eliciting more reasoning through chain-of-thought often fails to correct the error. Understanding what determines which reasoning signal ultimately controls a model's decision is therefore critical for building reliable LLMs, yet remains largely unexplained at the mechanistic level. Using the Heuristic Override Benchmark, we combine layer-wise workspace analysis, circuit tracing, and held-out causal interventions across model families. We uncover a consistent depth-dependent competition: models initially favor heuristic shortcuts, while successful reasoning depends on a late-layer reversal toward the constraint-consistent answer; activating this pathway rescues 34–60% of held-out failures across five models, while reverse interventions produce direction-specific degradation. Together, these results support a mechanistic account of shortcut reasoning as competition between heuristic-driven and constraint-sensitive computations, and identify a late corrective pathway whose activity predicts, and whose intervention causally controls, whether an initially favored shortcut is overridden. The recurrence of this pattern across models suggests a shared computational motif for resolving conflicts between salient heuristics and implicit constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.