TriSteer: Tri-Space Adaptive Steering for Large Audio Language Model Safety
Abstract
Large Audio Language Models (LALMs) have extended language-model interaction to the audio modality, while introducing new safety risks. We investigate how safety-relevant information is organized in audio-conditioned hidden representations and whether this structure can support inference-time control. A central challenge is to identify which internal semantic and acoustic evidence should trigger and shape an intervention. Our analysis reveals layer-dependent alignment between audio-native and text-only safety directions, while semantic and acoustic variations form reusable structures in the representation space. Based on these findings, we propose TriSteer, a training-free framework that constructs ordered contrast-defined measurement subspaces for semantic, decision-relevant acoustic, and decision-irrelevant acoustic variation. At each layer, TriSteer adaptively sets intervention strength from the current input's deviation relative to fixed benign reference boundaries, while acoustic attribution selectively refines the steering direction. Experiments across three LALMs and multiple safety benchmarks show that TriSteer improves the safety-utility trade-off without model updates. These results demonstrate that the geometry of internal safety representations provides a practical basis for adaptive inference-time control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.