acceptodds
Under review as a conference paper at ICLR 2027

Robustness Beyond Pixels for Vision Language Models

Abstract

Prompt learning adapts frozen vision–language models using a small number of trainable tokens, but adversarial prompt-learning methods typically defend only against image-space perturbations. We study a complementary white-box setting where an adversary jointly perturbs the input image and selected residual-stream activations, as may arise in shared-encoder or untrusted-adapter pipelines. This compositional threat model is challenging because pixel and feature perturbations interact through a common forward pass, while activation scales vary across transformer depth. We propose JPFAT (Joint Pixel–Feature Adversarial Prompt Tuning). JPFAT uses alternating projected gradient ascent over a fixed product threat set, Relative Activation Budgets to calibrate feature radii to layerwise activation scale, and entropy-regularized effort allocation to prioritize attack paths without changing their feasible radii. A vulnerability-adaptive consistency loss emphasizes samples with larger output and internal representation drift. Using fixed-budget, non-adaptive pixel, feature, sequential, and joint attacks for evaluation, JPFAT improves accuracy under the joint pixel–feature adversary and its individual components while maintaining competitive clean base-to-new generalization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.