Role-Consistent Attention Steering for Hierarchical Instruction Following
Abstract
Large language models often violate the instruction hierarchy by following lower-priority instructions that conflict with higher-priority ones. Attention steering can mitigate these violations by emphasizing higher-priority messages, but may also suppress lower-priority context needed for the task. Our analysis shows that this conflict–preservation trade-off depends strongly on head selection, with larger conflict benefits tending to incur larger preservation costs. We introduce **Ro**le-**C**onsistent **A**ttention **St**eering (**RoCAST**), a training-free method that selects a fixed subset of attention heads offline for role-based reweighting. RoCAST scores heads using conflict benefit across role-swapped instructions, role consistency, and preservation without conflict, with calibration requiring only 64 instruction pairs and 250 unlabeled paragraphs. Across five Llama and Qwen models spanning 3B–14B parameters, RoCAST improves IHEval conflict accuracy by 21.3–33.1 percentage points over native models and outperforms the evaluated inference-time baselines. It keeps aligned task execution within 1.3 points of native on every model, improves system-level instruction following, and incurs smaller capability losses than the baselines. Ablations show that role consistency and preservation improve both conflict and aligned performance over conflict-only selection. The same calibrated head sets transfer without recalibration to thinking-enabled inference and hierarchy-trained checkpoints, yielding further conflict gains while largely preserving aligned performance. Together, these results show that resolving instruction conflicts without sacrificing compatible behavior depends critically on which heads are steered, supporting offline head calibration as a lightweight complement to hierarchy-specific training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.