acceptodds
Under review as a conference paper at ICLR 2027

Trigger Appearance Is Not Enough: Evidence Routing for Diverse Backdoor Activation in Large Vision-Language Models

Abstract

Backdoor attacks against large vision-language models (LVLMs) take increasingly diverse forms. Backdoor triggers can be stealthy, adaptive, or distributed across modalities. Existing analyses built on trigger inversion or attention behavior rely on a specific assumption about what a trigger looks like and break down when the true trigger form does not match it, leaving little common ground for studying these otherwise disparate phenomena. We propose a unified perspective on backdoor activation: rather than asking what a trigger looks like, we examine whether an input causes the model to reach its output through an abnormal evidence route during autoregressive generation. This view treats diverse trigger mechanisms as instances of the same underlying phenomenon, an abnormal change in how evidence is routed toward the output. Our framework localizes a decision moment at which the committed token strongly depends on multimodal evidence, constructs an evidence support graph over visual patches, textual tokens, and their causally verified cross-modal interactions through counterfactual masking, and compares the resulting route against clean evidence output behavior estimated from a small calibration set, requiring no trigger inversion, no poisoned samples, and no prior knowledge of the target response. Across multiple LVLMs, datasets, and trigger types, this evidence route deviation remains a consistent signature of backdoor activation regardless of trigger form, and its use for detection consistently achieves AUROC and AUPR greater than 0.99.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.