acceptodds
Under review as a conference paper at ICLR 2027

GradTrace: Gradient-Neighborhood Filtering for Backdoor Defense in Multimodal Instruction Tuning

Abstract

Multimodal instruction tuning facilitates the adaptation of multimodal large language models (MLLMs) to downstream tasks but exposes them to backdoor poisoning attacks through potentially untrusted training data. Defending against such attacks requires identifying poisoned examples without trusted clean references while preserving useful task supervision. To address this challenge, we propose GradTrace, a reference-free data filtering method that exploits gradient-neighborhood structure to identify and remove suspicious training examples. Our key idea is to leverage local gradient consistency to identify suspicious samples and further assess their group-level separation before filtering. Specifically, GradTrace constructs compact representations of per-example adapter gradients, measures their directional agreement through neighborhood-based scoring, and applies a separation-aware decision rule to filter candidate groups only when sufficient separation is observed. Otherwise, all examples are retained to avoid unnecessary data removal. Extensive experiments across multiple MLLMs, datasets, and backdoor attack settings demonstrate that GradTrace effectively reduces attack success rates while preserving clean-task performance. Compared with existing defenses, GradTrace achieves competitive detection performance with substantially lower data cleaning time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.