acceptodds
Under review as a conference paper at ICLR 2027

ATLAS: Runtime Detection of LLM Backdoor Activation via Attention–Semantic Misalignment

Abstract

Backdoored large language models can behave normally on benign inputs while exhibiting attacker-controlled behavior only when a trigger is present. Dynamic backdoors are particularly challenging because their triggers can be expressed through latent input properties such as style or syntax rather than fixed lexical patterns, limiting the effectiveness of surface-level inspection. This motivates runtime detection based on model-internal behavior. However, internal activations naturally vary substantially across inputs, making it difficult to isolate changes caused specifically by backdoor activation. Despite this input-induced variability, we identify attention–semantic misalignment (ASM) as a systematic signature of activated executions, where attention becomes abnormally concentrated on tokens with low semantic relevance. Motivated by this observation, we present ATLAS (Attention Trace-based Latent Activation Surveillance), a white-box runtime detection framework for data-poisoning backdoors in LLMs. ATLAS characterizes normal ASM from trusted clean executions and aggregates reference-normalized deviations into an execution-level anomaly score, requiring neither trigger knowledge, poisoned examples, nor a clean model copy. Experiments across multiple LLM families, downstream tasks, and both static and dynamic backdoors show high detection rates with low false-positive rates. We further quantify the runtime cost of ATLAS, showing that it supports deployment-time monitoring with modest latency and memory overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.