CAVE: Counterfactual Auditing of Visual Evidence for Budgeted Local Supervision
Abstract
Medical images often have diagnosis labels but few local masks to supervise predicted evidence. With a limited annotation budget, we must select masks before their target regions are visible. Uncertainty and diversity identify ambiguous or underrepresented images but do not test whether predicted regions contribute to the known-class response. Expert-ROI criteria require the mask that the query will acquire. We introduce Counterfactual Auditing of Visual Evidence (CAVE), which compares known-class responses when a predicted region is retained or removed and measures reflection stability before ranking mask queries. Acquired masks supervise the shared learner. Under the original validation-selected five-seed profiles, CAVE has higher mean Dice budget-AUC than its matched control on BUSI, ISIC2018, and RSNA and the highest RSNA mean among tested selectors, although paired intervals span zero. Under a separate ten-seed protocol counting every queried mask, its mean Dice budget-AUC is lower than the matched control on all three datasets. The audit measures spatial evidence before annotation, but its current score does not reliably identify valuable masks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.