DriftWatch: Inference-Time Defense against Backdoored LLM Generation via Internal Anomaly Signals
Abstract
Backdoored large language models can behave normally on benign prompts while producing attacker-desired content under hidden activation conditions. Defending against such attacks in open-ended generation is challenging because malicious behavior can emerge in diverse forms during autoregressive decoding, and is not limited to a fixed label space. Existing defenses often rely on retraining, clean data, output-side probability signals, or repeated target-reference comparison, providing limited visibility into the internal dynamics of backdoor activation. We propose DriftWatch, an inference-time defense that monitors the model's own internal dynamics during decoding and performs localized token replacement when anomalous dynamics emerge. The proposed method computes anomaly evidence from the target model's hidden states and attention patterns, combining semantic drift, representation instability, and attention-trajectory signals into an adaptive anomaly score with an adaptive threshold. When a token is flagged, it is locally replaced using a reference model. Otherwise, the target model remains the primary generator. Experiments across diverse generative backdoor settings show that DriftWatch substantially reduces attacker-desired generations while preserving response quality. Its anomaly trajectories further provide generation-time observability, revealing when internal deviations emerge and how different signal families align with attacker-desired generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.