Probing and Steering Constraint Following in Multi-Turn Dialogue
Abstract
Why do language models stop satisfying an initial constraint as a dialogue grows? We construct a multi-turn evaluation set from IFEval-inspired instructions and Dolly questions: 23 constraint types, 230 sessions and 20 turns per session. Outcomes use repository-specific checks, not official IFEval scoring; some checks approximate rather than fully validate the requested constraint. Across Qwen2.5 3B, 7B, and 14B and Llama3 8B, we extract 18,400 pre-response states, excluding each current answer. Session-disjoint linear probes identify the constraint on 48.3–72.2% of subsequently checker-failed turns; all four family-adjusted intervals exceed the 4.35% uniform-label reference. This establishes decodability, not functional use. We compare concept activation vector (CAV) and instruction-directed attention steering in 108 common-cap cells totaling 496,800 response-condition evaluations. At unit strength, combined steering improves checker pass rates by 21.7–29.8 percentage points over baseline, with subadditive gains. Across all 36 model–strength combinations in the complete joint grid, adding CAV to attention both gains and loses successes; losses exceed gains in 24. Modest CAV steering adds value to weak attention, but no tested addition to stronger attention has a positive pointwise 95% interval excluding zero. Individual rescues therefore do not guarantee net joint benefit. The findings concern this exploratory evaluation protocol, not a demonstrated natural failure mechanism or semantic answer quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.