acceptodds
Under review as a conference paper at ICLR 2027

Post-Training Clinical-Signal Auditing of Medical Diagnostic Models

Abstract

In diagnostic imaging, specific findings vary in appearance with disease severity or clinical state, and these relationships, accumulated as clinical knowledge, provide an important basis for diagnostic interpretation. Diagnostic performance is a central axis for evaluating medical AI, but from diagnostic performance alone it can be difficult to determine whether models with similar performance respond differently to these changes. Using clinical ordering information withheld from training and final model selection as a reference, we assess how closely frozen-model responses follow the order of these changes within the same diagnostic label, and propose this assessment as a retraining-free post-training audit method. Across backbone–task combinations on four medical image diagnosis tasks (Liver US, Breast US, Knee X-ray, and Fundus), this degree of agreement ranged from to , and its relationship with diagnostic performance varied by task. In the three tasks with a coarse spatial scale prespecified for the relevant clinical findings, higher agreement was associated with better preservation of this ordering after controlled perturbations, and the association remained positive after adjustment for relevant covariates (). Because the audit method requires only frozen-model forward outputs on the model side, it can be applied to publicly released diagnostic checkpoints without access to their original training pipelines or internal representations. These results provide a complementary post-training criterion for assessing how closely model responses agree with the direction of change indicated by clinical knowledge, beyond what conventional diagnostic performance alone readily captures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.