DGTune: Discriminative Geometry Tuning for Noise-Affected Medical Vision–Language Models
Abstract
Pre-training on large-scale datasets followed by fine-tuning on downstream tasks has become a popular paradigm for medical vision–language models. However, medical descriptions and answers provide imperfect upstream supervision, raising a fundamental question: does clean downstream learning overcome the noise inherited from pre-training? We investigate this problem through a controlled family of medical vision–language models trained with different levels of supervision noise. Our experiments reveal persistent downstream performance deficits, including under distribution shift, despite clean adaptation labels. Through Fisher analysis, we link these persistent deficits to weakened class-discriminative structure in the adapted representations. Crucially, strengthening discrimination on observed samples alone does not ensure that the same discriminative relations generalize to new samples. Motivated by this distinction, we propose Discriminative Geometry Tuning (DGTune), a lightweight adaptation method that jointly strengthens weak Fisher directions and promotes their cross-sample generalization. Its two complementary regularizers enhance weak class separation and preserve the corresponding directions and class-contrast patterns across balanced, disjoint sample groups. Experiments on controlled noisy and public medical models show improved in-distribution and out-of-distribution generalization with DGTune.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.