Robustness Beyond Pixels for Vision Language Models
Abstract
Prompt learning adapts frozen vision–language models using a small number of trainable tokens, but adversarial prompt-learning methods typically defend only against image-space perturbations. We study a complementary white-box setting where an adversary jointly perturbs the input image and selected residual-stream activations, as may arise in shared-encoder or untrusted-adapter pipelines. This compositional threat model is challenging because pixel and feature perturbations interact through a common forward pass, while activation scales vary across transformer depth. We propose JPFAT (Joint Pixel–Feature Adversarial Prompt Tuning). JPFAT uses alternating projected gradient ascent over a fixed product threat set, Relative Activation Budgets to calibrate feature radii to layerwise activation scale, and entropy-regularized effort allocation to prioritize attack paths without changing their feasible radii. A vulnerability-adaptive consistency loss emphasizes samples with larger output and internal representation drift. Using fixed-budget, non-adaptive pixel, feature, sequential, and joint attacks for evaluation, JPFAT improves accuracy under the joint pixel–feature adversary and its individual components while maintaining competitive clean base-to-new generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.