Not All Gains Are Equal: Controlled Attribution in Medical Foundation Model Deployment
Abstract
Local adaptation can improve a medical foundation model's performance, but a lower post-adaptation error alone does not reveal whether the gain comes from correcting systematic output shift, using local supervision, or changing the reliability of downstream decisions. We introduce a controlled deployment attribution framework that separates calibration-recoverable improvement, the supervision-dependent value of released predictions, and the decision-level effect of predictive uncertainty. We instantiate the framework on PanEcho using 10,175 echocardiography sequences linked to 8,283 reports with seven continuous measurements. Before crediting a replacement head with the full improvement, we find that train-only affine calibration recovers 74.5% of the specified macro normalized-MAE gap between released predictions and a historical local MLP. Released predictions are most useful as an anchor under limited local supervision, while their advantage narrows as labels increase. Propagating patient-group prediction intervals through fixed measurement-derived rules changes the downstream risk–coverage trade-off: false assertions decrease from 10.52% at 100% assertion coverage to 2.47% at 49.40% coverage. In this controlled deployment setting, the results support attributing observed local gains to calibration, supervision, and decision-level effects rather than to a single adaptation gain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.