acceptodds
Under review as a conference paper at ICLR 2027

PathoDC-VQA: Constructing Difficulty-Controlled Pathology VQA Data from Sparse Captions

Abstract

Pathology vision–language models (VLMs) require domain-specific supervision for diagnostic reasoning, yet many public resources contain only information-sparse image–caption pairs. Naively converting these pairs into visual question answering (VQA) often produces caption paraphrases with little control over diagnostic difficulty. We introduce PathoDC, a three-stage pipeline that converts diagnostically eligible pairs into difficulty-controlled supervision. PathoDC first constructs structured Verified Diagnostic Knowledge (VDK) through automated generation and auditing. It then uses a VLM probe to organize candidates into single-image-answerable Easy cases and two evidence-sensitive hard tracks: knowledge-conditioned Hard-K and auxiliary-image-contrastive Hard-V. Finally, it generates track-conditioned rationales and applies an Easy→Hard-K→Hard-V training curriculum. This pipeline produces PathoDC-VDK with 6.3k records and PathoDC-VQA with 109k samples. Supervised fine-tuning of Qwen2.5-VL Instruct yields PathoDC-3B and PathoDC-7B. Using only about one-fifth of Patho-R1’s SFT data and no reinforcement learning, PathoDC achieves competitive performance on public pathology benchmarks. At the 7B scale, PathoDC-7B exceeds Patho-R1-7B on Quilt-VQA and performs comparably on PathVQA, while PathoDC-3B surpasses Patho-R1-3B on PathVQA. On the case-disjoint PathoDC-Bench, the two models improve their backbone weighted-average accuracy by up to 42.34 points and outperform the corresponding Patho-R1 models by 9.89 and 21.17 points, respectively. Although the gains are not uniform across all public VQA and PathMMU evaluations, the results show that evidence-aware difficulty control can improve pathology VLM supervision, with the strongest gains on diagnostic evaluations aligned with the proposed evidence taxonomy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.