Performance Without Evidence: A Placebo-Controlled Audit of Medical Foundation Model Adaptation
Abstract
In few-label adaptation of medical foundation models, AUC gain is routinely read as evidence that the model has learned the target pathology. We show that this inference is unsafe and argue that standard adapters frequently achieve high performance by covertly exploiting spurious artifacts rather than genuine lesions. Because purely observational metrics cannot expose this shortcut, we introduce a Placebo-Controlled Auditing framework to isolate true grounding by verifying that predictions degrade when the actual lesion is occluded, yet remain stable otherwise. The protocol first measures causal reliance with a paired counterfactual intervention that occludes the annotated lesion against a shape- and texture-matched control. It then validates that measurement with a lesion-free placebo: the identical comparison with no lesion present, which must stay silent for the intervention score to mean lesion reliance. Across eight public task beds and six frozen encoders per bed, standard adaptation raises AUC substantially, yet the trained adapter stays statistically indistinguishable from a prior-matched, gradient-free null in 44 of 48 cells. Because the capacity for grounded transfer cannot be reliably predicted by encoder provenance, it must be explicitly measured and enforced. To address this, we act via Counterfactual-Consistency Adaptation (CCA), which translates the validated causal contrast directly into a training constraint. CCA successfully forces reliance on genuine lesions, raising both causal evidence and external AUC on every CXR-domain encoder while outperforming six mask-matched baselines. Causal reliance is measurable, unpredictable, yet trainable once validated—positioning our measure–validate–act loop as the reporting standard for label-scarce clinical adaptation. The code is available: https://anonymous.4open.science/r/calm_official-08B4/README.mdhere.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.