acceptodds
Under review as a conference paper at ICLR 2027

From Risk Recognition to Safer Generation: Compete–Split–Fuse for Selective Representation Coordination

Abstract

Language models can correctly recognize a request's risks yet still produce unsafe content, exposing a gap between safety knowledge and behavior. We study this aware-but-unsafe (AUB) failure through paired risk judgments, token-level tracing, and activation interventions. The evidence supports a competition-based account: task-related representations can weaken safety's influence during generation, while autoregressive token-cache feedback can sustain the resulting trajectory. Motivated by this mechanism, we introduce Compete–Split–Fuse (CSF), an inference-time method for frozen language models. CSF detects prefixes at risk of unsafe execution, forms a task-release reference branch with a separate key-value cache, and selectively transfers only safety-private information back to the main branch before each token decision, leaving shared, task-private, and residual components unchanged at the fusion layer. Prediction consistency adaptively ends the extra computation. Across 8 checkpoints from 3 families, CSF reduces conditional unsafe assistance among pre-intervention knowledge-positive HarmBench requests from 25.4% to 10.6%, with broadly similar benign performance and a full-response serial-time proxy of 1.13× baseline. Mechanistic controls and ablations support predictive timing, task release, selective fusion, and persistent token-cache feedback as contributors to the observed gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.