Boldness: Evaluating Models under Label-Changing Edits via Diffusion Latent Optimization
Abstract
We define boldness as a model's ability to predict the new oracle label under valid nearby label-changing perturbations, complementing robustness under label-preserving perturbations. Evaluating boldness requires constructing natural, source-faithful samples whose oracle labels change. We characterize these requirements and formulate boldness-sample construction as a constrained optimization problem. To make this construction practical, we develop a plug-in diffusion method that optimizes the latent seed of a fixed pretrained conditional generator. The method combines target-class conditioning with source-label loss to search for label-changing edits on which the model retains the source label, while regularizing the latent seed around a source-inversion solution to promote source faithfulness. Target semantics and source faithfulness are validated before evaluation. Experiments on image classification and object detection reveal substantial boldness failures across the evaluated model families, including robustness-enhanced models, highlighting the need to evaluate boldness alongside clean accuracy and robustness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.