Towards Shared Flat Regions: Boosting Adversarial Transferability by Exponentially Tilting Curvature-Regularized Gradients
Abstract
Transfer-based adversarial attacks rely on accessible surrogate models to fool unseen targets, but often overfit surrogate-specific loss landscapes. Although flatness-aware attacks mitigate this issue, local flatness on a surrogate does not guarantee agreement with target models. We investigate neighborhood gradient stability as a surrogate-observable signal for constructing transferable attack directions, motivated by the hypothesis that effective perturbations exploit stable, high-loss regions shared across models. We propose Tilting Curvature Regularized Attack (TCRA), which combines response-guided finite-difference probing, parallel–orthogonal gradient decomposition, and exponential tilting. Local probes capture directional gradient changes, while tilting prioritizes candidate directions from high-loss neighborhoods. Our analysis bounds the deviation of the constructed direction and establishes sufficient conditions for target-loss ascent under model discrepancy, including momentum and projected sign updates. Experiments on ImageNet-compatible and CIFAR-10 benchmarks demonstrate improved transferability across standard, adversarially trained, and Transformer-based targets. TCRA also improves the evaluated gradient-based and input-transformation attacks without target queries or surrogate-parameter tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.