Reason Before You Generate: Latent Consistency Control for Radiology Report Generation
Abstract
Radiology report generation aims to automatically produce clinically coherent diagnostic reports from medical images. However, most existing RRG methods follow an open-loop autoregressive generation paradigm, where multimodal evidence is primarily used for token prediction, while the evolving report state is rarely re-evaluated against image-grounded semantics during decoding. To address this issue, we propose Latent Consistency Control (LCC), which establishes a shared cross-modal latent space and explicitly regulates semantic consistency throughout autoregressive generation. LCC first condenses visual and textual representations into compact semantic prototypes and performs iterative latent cross-modal reasoning to obtain a unified reasoning representation. Building on this latent space, we introduce a multi-step consistency-aware decoding mechanism that periodically evaluates the generated prefix against image-grounded semantics and conditionally injects a lightweight corrective signal when potential inconsistency is detected. In this way, LCC enables intermediate verification and corrective regulation rather than relying solely on sequence-level generation. Experiments on benchmark RRG datasets demonstrate consistent improvements in both natural language generation quality and clinical factual correctness, outperforming representative state-of-the-art methods on most metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.