acceptodds
Under review as a conference paper at ICLR 2027

MedCounterSeg: Benchmarking Unsupported Medical Queries and Learning to Defer through Concept-Based Visual Reasoning

Abstract

Recognizing a medical concept does not establish that a particular image supports it. A segmentation model must distinguish these judgments to avoid turning an unsupported query into a plausible mask. We introduce the Medical Query-Counterfactual Segmentation Benchmark (MedCounterSeg-Bench): 6,018 supported–unsupported query pairs across six imaging modalities, 42 concepts, and five clinically reviewed mismatch levels comprising ten subtypes. Each pair shares a target image, while each query receives valid positive examples of its requested concept. The benchmark separately measures concept localization, target-support decisions, and mask quality, with clinician involvement and iterative human quality control during construction. We further propose MedCounterSeg-RL, which learns to defer through concept-based visual reasoning. After positive-only concept–mask alignment, reinforcement learning jointly trains segmentation and deferral; correct deferral retains credit for locating the concept in a positive verification image. With Qwen3.5-4B, the method achieves the highest paired accuracy among evaluated methods on both the in-distribution Test split and unseen datasets, with segmentation quality close to ConceptSeg-R1 on in-distribution queries. The benchmark isolates fine-grained lesion and anatomical mismatches as the main remaining challenge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.