acceptodds
Under review as a conference paper at ICLR 2027

Trust What You See or What You Know? Evidence-Complementary Prototype Prompting for Biomedical Vision-Language Models

Abstract

Few-shot adaptation of biomedical vision–language models can draw on two kinds of evidence: text descriptions that know each class in general terms, and labeled images that show how the classes look in the target task. We observe that the two are reliable in opposite regimes: the text prior helps most when labeled images are scarce, while class prototypes built from the labeled images overtake it as the images accumulate. Most existing prompt learners, however, keep the influence of the text fixed no matter how many labels are available. A second problem appears when prototypes are simply added to a prompt trained with distillation: the distillation loss pulls their weight toward zero, because the text teacher has never seen the labeled images. We therefore propose Evidence-Complementary Prototype Prompting (ECPP), which lets each kind of evidence contribute where it is reliable. Concretely, (1) an evidence handoff gradually shifts trust from the text prior to a prototype classifier as more labeled images become available, and (2) text-path distillation keeps the teacher's guidance on the text branch, so that the prototype weight is learned from the labels alone; we show analytically why this placement matters. ECPP adds only one trainable scalar to a frozen BiomedCLIP prompt learner. On eleven biomedical datasets spanning nine imaging modalities, ECPP improves over the state-of-the-art BiomedCoOp at every support size (with gains of up to 5.34% on breast ultrasound) and achieves the best base-to-novel harmonic mean among the compared methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.