A quoted answer overrides the rule that designates it: what an instruction-compliance score measures depends on how the rule names its answer
Abstract
Conflict-only instruction-compliance scores ask whether a verifier endorses the rule-designated answer when the facts favour another. Their contrast, , subtracts correctness sensitivity from compliance sensitivity; without aligned items, it cannot identify the two separately. On four fixed instruction-tuned checkpoints, holding label indirection fixed and printing the permitted answer changes conflict-cell error by , with substantial variation across checkpoints. We then separate the rule's label-based designation from the answer it quotes. In a prospectively registered form where the two disagree, and with the factual question removed and both options replaced by invented words, accuracy against the label binding is 0.103, versus 0.955 when they agree. The quoted option is endorsed at 0.806, versus 0.012 for the label-designated one. The label binding remains usable without a quotation (balanced accuracy 0.922); thus factual-answer recall cannot explain this override of the label-based designation in these constructed rules. For property-based rules, balanced accuracy without a factual question is 0.526 [0.518, 0.535]: below the registered competence floor, but not at chance under its registered criterion. The verdict is indeterminate, so their conflict-cell cost cannot be isolated as a rule-form effect. An independent equal-budget comparison also did not establish a selection advantage for the four-cell audit. These empirical findings concern this fixed small-model roster and two-option tasks; the identification problem applies to the score's definition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.