acceptodds
Under review as a conference paper at ICLR 2027

VEC-CT: Visual Cue-Guided Error-Corrective Clinical Reasoning for Lung CT Report Generation

Abstract

Multimodal large language models (MLLMs) provide strong capabilities in visual perception and reasoning, offering a promising approach to medical image understanding and diagnostic report generation. However, accurate interpretation of lung CT scans remains challenging due to subtle lesion characteristics, complex anatomical structures, and strong dependencies among clinical findings. Moreover, inaccurate visual observations in early clinical sub-tasks may propagate and accumulate through subsequent reasoning steps, ultimately reducing the accuracy and clinical consistency of the generated reports. To address these challenges, we construct the LungCT-CoT dataset, which decomposes lung CT report generation into six clinically ordered and interconnected sub-tasks to reflect the clinical diagnostic workflow. Based on this dataset, we propose VEC-CT, a visual cue-guided error-corrective clinical reasoning framework. The model first performs structured reasoning over the predefined clinical sub-tasks and generates an initial diagnostic report. The Error Tracer then identifies the earliest erroneous sub-task in the reasoning process, after which the Visual Cue Extractor extracts task-specific visual cues relevant to the identified sub-task from the CT image. Guided by these visual cues, the model revisits the erroneous sub-task and corrects the subsequent reasoning process to refine the final diagnostic report. Extensive experiments show that VEC-CT consistently outperforms representative methods on both conventional text generation metrics and clinically oriented semantic evaluation metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.