LLM-DE: Jailbreaking LLM-based Malware Detection via Prompt Injection and Safety Alignment Exploitation
Abstract
Malware detection is entering a third generation in which Large Language Models (LLMs) reason over textual renderings of program features, promising contextual understanding that classical byte- and feature-space classifiers lack. This shift relocates the attack surface: the verdict is computed over text derived from the file, a channel that is fully attacker-controllable. We present LLM-DE, a systematic study of adversarial evasion against LLM-based malware detectors. LLM-DE formalizes four attack categories—prompt injection via feature descriptions (PIPR-A), safety-alignment exploitation (SAE-A), parsing-error induction (PEI-A), and semantic camouflage (SC-A)—each targeting a distinct stage of the LLM detection pipeline under a strict functionality-preserving, feature-addition-only constraint. We evaluate end-to-end against detectors built on six backbone models spanning three orders of magnitude. The evaluation uncovers a capability prerequisite: sub-threshold backbones issue input-invariant verdicts, so the detector—not the attack—fails first. On the smallest discriminating backbone (Phi-3-mini, measured), every payload-carrying construction flips every measured malicious sample to confidently benign (residual malicious probability for the unified optimized attack), and effectiveness varies across backbones in a per-channel manner—injection is near-universal, while alignment exploitation, semantic camouflage, and truncation scale with alignment strength, semantic priors, and context budget, respectively. A refusal-gate mechanism analysis—grounded in how aligned LLMs attribute harmful intent rather than harmful keywords—explains the naive variants' failures and yields optimized constructions whose measured outcomes exceed their disclosed projections. Against four baseline attack strategies, single-channel baselines (payload-only, suffix-only, or feature-space-only) evade at most one detector family, while the joint LLM-DE construction evades both the LLM detector and a trained classical DREBIN-SVM (99.7% ESR) simultaneously. Our results indicate that moving malware detection to LLMs does not remove adversarial vulnerability; it re-opens it along four channels that no single backbone property closes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.