acceptodds
Under review as a conference paper at ICLR 2027

PROBE: Automated Spurious Cue Discovery and Cross-Architecture Shortcut Scoring

Abstract

Image classifiers often rely on shortcuts, yet systematically discovering and comparing such dependencies across classifiers remains challenging. We introduce PROBE, a fully automated two-stage auditing pipeline requiring no predefined attributes, group labels, or retraining. The first stage, Discover, names the visual cues a classifier relies on for each class and tests each cue through a targeted counterfactual edit, measuring the resulting change in classifier confidence to separate object-intrinsic from spurious cues. The second stage, Score, turns the discovered cues into rendered probe images that show the object with its context, the context alone, the object alone, or the object under stacked confuser cues, and reports object-only accuracy, shortcut-stress accuracy, and confidence shift for each classifier. Across five ImageNet classifiers evaluated on a shared candidate set, 86% of shortcut cues that pass each classifier's own confidence threshold are flagged by at least two classifiers and 31% by all five. Probes that show the context alone rarely recover the label after filtering detected render leakage. The discovered catalog recovers $4% of human-annotated reference cues and achieves higher measured per-phrase precision against Salient ImageNet labels than three annotation-free baselines. Debiasing a classifier moves the audit's scores in the expected direction, while an untrained classifier yields no substantive findings, indicating that the scores respond to classifier behavior rather than being determined solely by the tools that construct the probes. The pipeline scales to thousand-class and open-vocabulary settings, where the extreme pairwise differences on the accuracy-based axes reach statistical significance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.