ClinKD: Cross-Modal Distillation for Fine-Grained Medical VQA
Abstract
Fine-grained medical visual question answering requires a model to identify findings, ground them spatially, and generate clinically relevant text. We investigate two potential weaknesses in adapting a general-purpose multimodal model to this setting: contiguous position indices do not distinguish image-text boundaries, and uniform distillation can overemphasize uncertain teacher targets. ClinKD addresses these issues with Modality-aware Cross-modal Gap Rotary Position Embedding (MCG-RoPE) and confidence-margin curriculum distillation. MCG-RoPE inserts explicit boundary gaps while preserving the native spatial grid, whereas the distillation objective weights teacher-forced token distributions by confidence and top-two probability margin and gradually relaxes its admission threshold. On Med-GRIT-Test30k, the full pipeline obtains 67.51, 82.35, 70.56, and 65.69 on visual grounding, referring object classification, referring captioning, and medical image analysis, respectively, averaging 71.53. This is 14.87 points above the reported BiRD baseline, which uses a different backbone and is therefore not architecture-matched; on LLaVA-Med-ga0.2k, mBMR increases from 21.06 to 23.54. Single-component experiments show gains from MCG-RoPE and distillation separately, although the archived results do not isolate every interaction in the full pipeline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.