acceptodds
Under review as a conference paper at ICLR 2027

CONFIDENCE IS NOT EVIDENCE: PROBING CLASSIFIER SENSITIVITY WITH STEERABLE LENS

Abstract

Moving an image along a classifier's gradient is a common way to probe what change would carry it across a decision boundary, but crossing that boundary does not by itself establish a recognizable change toward the target class. We introduce Steerable Lens, a framework for investigating classifier sensitivity through complex steerable-pyramid (CSP) interventions, with pixel and Fourier-phase edits as comparators. At matched image-change budgets, we measure three quantities separately: the optimized classifier's confidence, transfer to independent classifiers, and compatibility with a smooth image transformation. These measurements yield different method rankings. A pixel perturbation can drive the optimized classifier to near-certain target confidence, above 96% on MNIST digits, yet independent classifiers rarely agree on the target, and the edit is poorly captured by a smooth deformation. CSP edits align closely with a smooth warp-plus-gain model while classifier confidence stays low. On CelebA the structural ordering between representations depends on the reconstruction model. In a pilot human study, CSP edits receive more ratings of clear attribute change and no visible corruption than pixel or Fourier-phase edits. Frozen band probes further reveal fine-scale sensitivity in the examined driver classifiers, while larger budgets reduce common attainment across bands and controls. These results position CSP as a controlled diagnostic representation and motivate evaluating optimized confidence, classifier transfer, and smooth reconstruction separately, without treating any one as certification of recognizable target-class change.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.