Auditing Image-Induced Errors in Vision-Language Models
Abstract
Small image perturbations can change vision-language model (VLM) answers, but pixel constraints and output changes alone do not establish whether an internal signal identifies image-induced errors. Such a signal must distinguish harmful perturbation effects from legitimate visual dependence; its usefulness against ordinary attacks must also be separated from its vulnerability to adaptive attacks. We study these questions through Boundary Violation Rate (BVR), which compares visual and trusted-text value-zero sensitivity. Attack optimization and final detection are specified separately: a cheaper optimization surrogate does not redefine the detector used to assess evasion. A matched-estimator analysis motivates SRBV-PGD as an exposed positive control, while SRBV-stealth suppresses reference-answer exposure during output disruption. Existing LLaVA-1.5/POPE experiments show 100% normalized yes/no reversals across 360 evaluations with near-zero group-only BVR. This establishes a limitation of that reference-based diagnostic, not evasion of the hybrid detector. The current evidence motivates matched actual-answer validation and fixed-detector evaluation, rather than establishing universal error attribution or reliable deployment-time detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.