acceptodds
Under review as a conference paper at ICLR 2027

SafeDepth: Safety-Aware Token-Level Adaptive Computation

Abstract

Recent studies suggest that not every token needs to pass through all Transformer layers, motivating token-level adaptive models that selectively skip layers to re- duce computation. Our experiments show that these execution choices also affect safety: existing token-level adaptive reasoning models exhibit higher harmful- response rates than their reference backbones. The safety effects depend on which layers are skipped and whether skipping occurs during prompt processing or an- swer generation. We further find that recognizing harmful requests and producing refusals depend on different layer-wise computations. We introduce SafeDepth, a lightweight, plug-in framework that uses selective layer execution to improve both safety and efficiency. SafeDepth learns to retain computations that support safe responses and bypass those that contribute to unsafe generation. A router se- lects layer execution based on token representations, layer position, and inference phase, while an adapter supports the skipped paths. We jointly train these modules to reduce computation and harmful responses while preserving performance on benign tasks, targeting a better safety–efficiency trade-off. The pretrained back- bone remains frozen throughout, requiring no additional pretraining. Experimen- tal results on Llama-3-8B-Instruct show that SafeDepth reduces computation rela- tive to full-depth inference while largely preserving task performance. Compared with FlexiDepth, it lowers unsafe-response rates across all five harmful-request benchmarks, with the largest absolute decrease on HarmBench-HJ (from 55.44% to 25.25%), while reducing the XSTest false-refusal rate from 12.0% to 4.0%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.