CONFIDENCE IS NOT EVIDENCE: PROBING CLASSIFIER SENSITIVITY WITH STEERABLE LENS
Abstract
Moving an image along a classifier's gradient is a common way to probe what change would carry it across a decision boundary, but crossing that boundary does not by itself establish a recognizable change toward the target class. We introduce Steerable Lens, a framework for investigating classifier sensitivity through complex steerable-pyramid (CSP) interventions, with pixel and Fourier-phase edits as comparators. At matched image-change budgets, we measure three quantities separately: the optimized classifier's confidence, transfer to independent classifiers, and compatibility with a smooth image transformation. These measurements yield different method rankings. A pixel perturbation can drive the optimized classifier to near-certain target confidence, above 96% on MNIST digits, yet independent classifiers rarely agree on the target, and the edit is poorly captured by a smooth deformation. CSP edits align closely with a smooth warp-plus-gain model while classifier confidence stays low. On CelebA the structural ordering between representations depends on the reconstruction model. In a pilot human study, CSP edits receive more ratings of clear attribute change and no visible corruption than pixel or Fourier-phase edits. Frozen band probes further reveal fine-scale sensitivity in the examined driver classifiers, while larger budgets reduce common attainment across bands and controls. These results position CSP as a controlled diagnostic representation and motivate evaluating optimized confidence, classifier transfer, and smooth reconstruction separately, without treating any one as certification of recognizable target-class change.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.