Same Conversation, Different Actions: Phase-Conditioned Causal Steering for Multi-Turn LLM Safety
Abstract
In multi-turn LLM dialogue, the same internal correction can improve one answer but impair another, even when both contexts require the same response behavior. We find that these differences are reproducible and that relative dialogue trajectories help select useful corrections. In this paper, we introduce PhaseSteer, which learns a shared correction family and its allocation from complete-response outcomes with the backbone frozen. A paired objective rewards appropriate answers across clarification and refusal boundaries, with separate constraints on quality loss and added harm. On our paired benchmark, PhaseSteer improves assistance-to-clarification and clarification-to-refusal joint success over development-selected comparators by an average of 4.1 and 4.5 percentage points across four backbones. Controlled analyses show improved allocation among the same learned corrections, while interactive tests show better boundary handling under generated histories. These gains are achieved with one action sampled per answer, without scoring candidate responses at inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.