TRACE: Temporal Risk Assessment and Concept Editing for Safer Language Generation
Abstract
Large language models (LLMs) can produce harmful content while generating a response. A useful safeguard should not only detect this risk during generation but also intervene while leaving safe responses largely unchanged and provide an auditable record of the internal patterns supporting its decisions. We propose TRACE (Temporal Risk Assessment and Concept Editing), a framework that tracks harmful-response risk during generation, selectively edits token activations, and leaves an auditable trace of its detection and steering decisions. TRACE learns a codebook of recurring activation patterns, with each codebook vector representing a latent concept. At inference time, it assigns each token activation to the nearest codebook vector to obtain the corresponding concept. A gated recurrent unit (GRU) processes the activations and their assigned concept features to predict completed-response harmfulness and a monotone cumulative streaming risk. During steering, activations assigned to harmful concepts are steered along directions defined by nearby benign targets. An external LLM labels the learned concepts post hoc, making detection sequences and harmful-to-benign edit records interpretable. Experiments across three generators and the WildGuard and S-Eval benchmarks show that TRACE improves streaming F1 over Kelp in five of six comparisons while remaining competitive on completed-response detection. TRACE's selective steering reduces Harmbench attack success with little change in refusal on originally safe responses. The post-hoc audit reveals recurring semantic structure in the learned concepts and distinct traces associated with harmful detection, safe generation, and activation edits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.