acceptodds
Under review as a conference paper at ICLR 2027

Boldness: Evaluating Models under Label-Changing Edits via Diffusion Latent Optimization

Abstract

We define boldness as a model's ability to predict the new oracle label under valid nearby label-changing perturbations, complementing robustness under label-preserving perturbations. Evaluating boldness requires constructing natural, source-faithful samples whose oracle labels change. We characterize these requirements and formulate boldness-sample construction as a constrained optimization problem. To make this construction practical, we develop a plug-in diffusion method that optimizes the latent seed of a fixed pretrained conditional generator. The method combines target-class conditioning with source-label loss to search for label-changing edits on which the model retains the source label, while regularizing the latent seed around a source-inversion solution to promote source faithfulness. Target semantics and source faithfulness are validated before evaluation. Experiments on image classification and object detection reveal substantial boldness failures across the evaluated model families, including robustness-enhanced models, highlighting the need to evaluate boldness alongside clean accuracy and robustness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.