LatentPrism: Adaptive Hidden-State Decoupling for Robust and Near-Lossless LLM Backdoor Defense
Abstract
Large Language Models (LLMs) remain vulnerable to backdoor attacks that embed covert trigger-to-target associations during fine-tuning. Existing defense paradigms—including data-side filtering and model-side post-training repair—predominantly inspect internal hidden states to identify backdoor anomalies. However, our empirical analysis reveals that relying on a single fixed metric or layer suffers from substantial performance fluctuations across network depths and attack paradigms. We identify this limitation as the layer-feature specificity trap. To overcome this trap, inspired by the spectral decomposition of a physical prism, we decompose high-dimensional hidden states into six complementary statistical features that characterize backdoor-induced anomalies from distinct geometric perspectives. Guided by this multi-perspective metric, we propose LatentPrism, an unsupervised backdoor defense framework that automatically selects the optimal layer-feature configuration without requiring clean reference models or labeled data. With this configuration, we reliably distinguish poisoned samples from clean ones via unsupervised clustering. Based on this sample partition, LatentPrism executes a dual-track purification process: Track 1 fine-tunes the model via sequence-level unlikelihood loss to suppress malicious generation pathways on poisoned samples, while Track 2 synthesizes benign pseudo-targets and fine-tunes a lightweight LoRA adapter to restore benign utility. Across three open-weight LLMs, five backdoor attack paradigms, and two security-critical downstream tasks, LatentPrism reduces average ASR from 64.7–99.6% down to at most 2.3% (achieving exact 0.00% ASR on 19 of 30 settings) while maintaining competitive utility across eight standard benchmarks under a strict 500-sample calibration budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.