FusionBreak: Attacking Multimodal Fusion While Preserving Unimodal Decisions
Abstract
Multimodal models combine information from multiple modalities, but correct unimodal predictions do not necessarily guarantee a correct fused prediction. Conventional adversarial evaluations typically focus on whether the final multimodal prediction becomes incorrect without distinguishing errors that emerge at the fused output from those already present in individual modalities. We study a stricter condition in which the unimodal classifiers remain correct while the fused model becomes incorrect. We introduce FusionBreak, an adversarial attack that preserves unimodal decisions while causing the fused model to misclassify, and define the Fusion-Only Success Rate (FOSR) to measure this behavior. Experiments across two datasets and four fusion schemes show that fusion-only failures occur across different fusion designs. In contrast, standard PGD can achieve a high fused attack success rate with 0% FOSR, showing that conventional attack success does not distinguish whether the fused model fails while the unimodal predictions remain correct. Ablation experiments further demonstrate the importance of preserving unimodal decisions. We also evaluate a detector-aware variant of FusionBreak that adds detector-evasion terms to the attack objective. This reduces detection by several generic adversarial detectors while largely preserving the effectiveness of the attack.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.