acceptodds
Under review as a conference paper at ICLR 2027

Layer-Wise Spatial Evidence for Medical Visual Grounding

Abstract

Medical visual grounding requires identifying the intended lesion or anatomy and delineating its precise extent under limited pixel coverage, ambiguous boundaries, and cross-modal domain shift. Existing segmentation-token methods typically select a Transformer layer based on hidden states before constructing its spatial map. However, an isolated hidden state lacks the spatial context to reveal where its eventual spatial map will respond, resulting in suboptimal layer composition and background contamination. We introduce LaSE, the first medical grounding framework to adopt a map-first principle: it evaluates spatial evidence prior to layer fusion. LaSE comprises two novel components: (1) Layer-wise Token Matching (LTM), which constructs a candidate spatial map at every Transformer layer by matching the generated segmentation state with image tokens, exposing layer-wise spatial responses without pre-selecting a layer; and (2) Medical Evidence-Guided Fusion (MEGF), which pools features from the medical regions indicated by each candidate map in a shared frozen MedSAM2 representation and adaptively predicts case-dependent weights, thereby suppressing candidates that point to spurious anatomies. We further propose Lightweight Mask Refinement (LMR), which refines the decoded mask with a lightweight convolutional pyramid and requires only a single MedSAM2 decoder call. Extensive experiments on a four-source (Source-4), ten-target (OOD-10) medical segmentation benchmark demonstrate the effectiveness of LaSE: it achieves 90.20% Source-4 and 84.14% OOD-10 DSC, surpassing the state-of-the-art MedCLIPSeg by 1.09 and 5.12 points on Source-4 and OOD-10, respectively, while improving the pre-decoder map Soft Dice from 54.55% to 59.92%. Code is available at https://anonymous.4open.science/r/LaSE-22.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.