acceptodds
Under review as a conference paper at ICLR 2027

DriftWatch: Inference-Time Defense against Backdoored LLM Generation via Internal Anomaly Signals

Abstract

Backdoored large language models can behave normally on benign prompts while producing attacker-desired content under hidden activation conditions. Defending against such attacks in open-ended generation is challenging because malicious behavior can emerge in diverse forms during autoregressive decoding, and is not limited to a fixed label space. Existing defenses often rely on retraining, clean data, output-side probability signals, or repeated target-reference comparison, providing limited visibility into the internal dynamics of backdoor activation. We propose DriftWatch, an inference-time defense that monitors the model's own internal dynamics during decoding and performs localized token replacement when anomalous dynamics emerge. The proposed method computes anomaly evidence from the target model's hidden states and attention patterns, combining semantic drift, representation instability, and attention-trajectory signals into an adaptive anomaly score with an adaptive threshold. When a token is flagged, it is locally replaced using a reference model. Otherwise, the target model remains the primary generator. Experiments across diverse generative backdoor settings show that DriftWatch substantially reduces attacker-desired generations while preserving response quality. Its anomaly trajectories further provide generation-time observability, revealing when internal deviations emerge and how different signal families align with attacker-desired generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.