Placement, Not Quantity: How Transformers Lock Into the Wrong Rule, and Why the Repair Signal Is Invisible to the Model
Abstract
A model trained under ambiguous supervision can lock into a wrong rule: a shortcut that gives the same labels as the intended rule on the training distribution, but different labels outside it. We ask what repairs such a model, and we find that the answer is placement, not quantity. A few counterexamples repair the lock-in only when they are placed to cover the region where the two rules disagree. How many counterexamples we add, how uncertain the model is about them, and how diverse they are do not matter. We show this in a controlled three-valued-logic testbed, where an intended rule and an inductively simpler shortcut agree on the training distribution by construction, so any preference of the model is inductive. Across three lock-in families, selecting counterexamples that cover the disagreement region beats selecting the same number at random or by diversity, with paired-bootstrap CIs that exclude zero. A direct-sampling intervention shows that region coverage is causally necessary at every budget we test: selection without coverage never closes the gap, even at four times the minimal pair. Eight placed counterexamples outperform an unplaced dose of 3200 on the held-out disagreement-region test, because off-target examples dilute the few that cover the region. Lock-in also follows a directional regularity: in all four families the model commits to the inductively simpler rule, whichever rule we declare correct, and re-scoring one family's weights under a flipped ground truth confirms the symmetry. Finally, in this testbed the repairing counterexamples cannot be found in advance. The signal is invisible to the model's confidence, to value ranking, and to ensemble disagreement: every locked seed commits to the same rule, so a committee is unanimous exactly where repair is needed (ranking AUC 0.500), and the model is confidently wrong where it errs. On a disagreement region of this kind, diagnosis must be interactive. A CelebA demonstration carries the placement signal to natural data: greedy failure-region selection beats budget-matched random with a paired CI excluding zero, on ten seeds pooled and on the five pre-registered extension seeds alone. There the failure region is one-sided and the locked model's own loss also finds it, so the invariant is the coverage requirement, not the selector.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.