Zero-Shot Context-Aware Medical Case Retrieval via Token-Set to Token-Set Similarity
Abstract
In healthcare, model-predicted outcome labels are, in practice, unhelpful, especially when clinicians are concerned about the uncertainty of these predictions. In contrast, references to past similar cases provide valuable guidance. Therefore, in this work, when presented with a new patient, our objective is to retrieve similar past medical cases to support treatment planning within a specific clinical context of interest. Previous methods rely on semantic searches and content-based image retrieval using vector databases of EHR and medical image embeddings. Although efficient, these approaches often lack contextual relevance, leading to lower retrieval performance. Recent multimodal foundation models can enable more effective context-aware cross-modal retrieval. While these models provide deeper context and richer embeddings, they are computationally expensive. Considering this, we propose a two-stage retrieval system that combines the efficiency of vector-based searches with the contextual depth of token embeddings provided by foundation models without further training. More importantly, we propose a token-set to token-set similarity for context-aware retrieval using rich token embeddings. Our extended experiments demonstrate the efficiency contributed by Stage 1 and the effectiveness contributed by the context-aware re-ranking in Stage 2.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.