acceptodds
Under review as a conference paper at ICLR 2027

MiragePoison:Few-shot Poisoning Attacks on Vision-Language Models for AIGC Detection

Abstract

Vision-Language Models (VLMs) have demonstrated strong multimodal understanding capabilities and are increasingly adopted for AI-generated content (AIGC) detection. However, their reliability in such security-sensitive scenarios, particularly under data poisoning during fine-tuning, remains underexplored. In this work, we systematically evaluate mainstream VLMs for AIGC detection and investigate their susceptibility to fine-tuning-based poisoning. We propose MiragePoison, a few-shot poisoning framework that constructs poisoned samples visually consistent with AIGC content while shifting their latent representations toward those of real images. Fine-tuning with only a small number of such samples systematically induces real images to be misclassified as AI-generated. On FakeVLM, MiragePoison achieves a 98.2% R2F-ASR with only 100 poisoned images, corresponding to merely 0.00000143% of the model parameter scale, demonstrating that minimal malicious supervision can substantially alter authenticity judgments. We further characterize poisoning-budget effects and compare the sensitivity of the Vision Encoder, Projector, and LLM. These findings show that limited malicious training data can significantly compromise VLM-based AIGC detectors and expose structural risks in their fine-tuning process.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.