acceptodds
Under review as a conference paper at ICLR 2027

Gradual Adversarial Transition Alignment for Robust Vision-Language Models

Abstract

Vision-language models such as CLIP exhibit strong zero-shot generalization but remain highly vulnerable to adversarial perturbations. Most existing adversarial fine-tuning methods primarily constrain the clean/reference state and the final adversarial state, leaving the progressive semantic drift between them unsupervised. A natural remedy is to supervise the intermediate states of iterative attacks, but this requires an additional forward–backward pass for every state, and real trajectories carry attack-specific local fluctuations. We therefore investigate whether attack trajectories exhibit a compact geometric structure that can be efficiently approximated. Our analysis reveals that, despite substantial local fluctuations, the cumulative representation displacement follows a stable dominant direction across attacks and perturbation settings. This finding motivates an endpoint-induced surrogate path that approximates the dominant clean-to-adversarial transition using only the two endpoints. Inspired by gradual domain adaptation, we discretize this path into a sequence of ordered intermediate states and enforce KL-based prediction consistency between adjacent states. The resulting Gradual Adversarial Transition Alignment (GATA) replaces a coarse endpoint constraint with a sequence of locally smaller predictive transitions, providing efficient transition-level supervision without explicitly modeling the full attack trajectory. Extensive experiments across diverse attacks, perturbation strengths, downstream datasets, and vision-language backbones demonstrate improved adversarial robustness while preserving a favorable clean–robust trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.