Differentiable Boundary Priors for Bayesian Ray Search in Hard-Label Attacks
Abstract
Hard-label attacks, one of the most challenging settings in black-box adversarial attacks, rely solely on top-1 label feedback and thus suffer from extremely high query complexity. Existing approaches, such as OPT, formulate hard-label attacks as a zeroth-order ray search problem. Subsequent methods build on this formulation and seek better query efficiency through refined gradient estimation or surrogate priors, yet their improvements remain limited under the extremely sparse feedback of the hard-label setting. In this paper, we propose B-OPT, a Bayesian optimization framework that improves hard-label ray search by optimizing a reparameterized objective under extremely limited feedback. We further propose PB-OPT, which goes beyond local gradient priors by incorporating a global functional prior into the optimization process. To enable this in the hard-label setting, we design a decoupled forward–backward mechanism: the forward pass estimates the surrogate boundary distance via parallel queries, while the backward pass approximates its gradient using a substitute loss. Based on this design, PB-OPT incorporates the surrogate objective as a functional prior in the Gaussian-process mean, allowing Bayesian optimization to refine its posterior estimate using both the surrogate prior and observed target evaluations. We derive a regret bound showing that informative functional priors tighten optimization bounds. Extensive experiments on image classifiers and vision-language models demonstrate that our method significantly improves query efficiency and outperforms 16 state-of-the-art hard-label attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.