Quality Hijacking: Semantic-Aware Typographic Attacks on Image Quality Assessment
Abstract
Vision-language model-based image quality assessment methods (VLM-IQA) have shown strong capability in evaluating complex and out-of-distribution images, yet remain vulnerable to typographic attacks. Attacking VLM-IQA requires adversarial text to effectively manipulate quality-related reasoning while remaining visually consistent with the underlying image. Existing typographic attacks typically rely on generic adversarial texts (ATs) and heuristic placement strategies, overlooking task-related factors and local image degradation, which limits both attack effectiveness and visual plausibility. To address these limitations, we propose Q-Hack, a black-box framework for typographic attacks against VLM-IQA models. Q-Hack adopts a coarse-to-fine strategy to jointly optimize AT content and placement. In the coarse stage, it explores three IQA-related attack factors to generate candidate texts, and naturally places them in suitable regions based on scene semantics, local degradation, and depth. In the fine stage, it exploits target-model feedback to distill effective semantic patterns and refine both text content and placement for stronger attacks. We compare Q-Hack with seven state-of-the-art baselines on three widely used IQA datasets, where it demonstrates clear advantages across attack effectiveness, perceptual fidelity, and scene consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.