PREDICTABLE BUT UNSTABLE : PROBING CONSTRAINT VIOLATIONS IN MULTI-TURN DIALOGUE
Abstract
In single-turn settings, linear probes over attention-head activations can predict whether a language model will violate an explicit constraint before generation. We extend this framework to multi-turn dialogue and find that mean AP does not systematically decline, yet probes become notably less stable when rebuilt from different development partitions—a pattern we term predictable but unstable. We identify two structural factors that may contribute: the first-violation task induces a survival structure that progressively thins supervision at later turns, even from an initial pool of 2,400 dialogues per configuration, and the relative construction stability of fixed attention-head banks varies across turns, so that a bank relatively stable at one turn may be less stable at another. These factors make probe construction sensitive to the turn composition of a development partition. We propose Frequency-Consensus head selection, which modifies only the aggregation rule while leaving the rest of the probing pipeline unchanged. On three 7B models, Frequency-Consensus improves model-averaged AP by 1.9 points over retaining all candidate heads while reducing cross-run SD by 37.5% and range by 34.8%, with the pattern extending to the evaluated 13–14B models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.