AFOC: ADAPTIVE FREQUENCY-ORTHOGONAL COUNTERACTION FOR ROBUST CLIP INFERENCE
Abstract
Zero-shot CLIP inference is vulnerable to adversarial perturbations, yet many defenses require training, labels, or prompt optimization. We study the more constrained setting of frozen encoders, unchanged prompts, and no training data. We propose Adaptive Frequency-Orthogonal Counteraction (AFOC), a training-free defense that estimates input risk from local embedding instability and sensitivity to frequency attenuation. A unified risk score calibrates orthogonal exploration, trajectory momentum, Fourier-domain update shaping, and a soft final gate, allowing conservative corrections for stable inputs and stronger corrections for risky inputs. Our local analysis formalizes clean-drift control and characterizes the directional and spectral effects of the update; this analysis is mechanistic rather than a robustness certificate. An embedding-geometry diagnostic further shows that PGD can disrupt zero-shot alignment while preserving local visual neighborhoods, whereas AFOC recovers decisions by rearranging embeddings relative to fixed text anchors. Across CIFAR-10, CIFAR-100, and ImageNet, we evaluate both post-hoc defense against attacks on the base classifier and defense-aware EOT-PGD attacks. On CIFAR-100 with CLIP ViT-B/32, AFOC attains 30.18% accuracy under adaptive EOT-PGD and a descriptive average of 39.22% across the four reported attacks, while retaining 56.55% clean accuracy. On ImageNet, it attains 29.53% adaptive accuracy with a 0.11-percentage-point drop in clean accuracy. These results demonstrate a competitive fixed-prompt operating point and motivate broader end-to-end adaptive evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.