acceptodds
Under review as a conference paper at ICLR 2027

VESP: VISUAL EVIDENCE and semantic consistency preservation IN AUTOREGRESSIVE MRI REPORT GENERATION

Abstract

Automated radiology report generation from 3D MRI remains challenging because visual evidence must remain associated with the appropriate clinical findings and influential throughout autoregressive generation. Existing approaches typically introduce visual information during encoding, but may suffer from insufficient field-level evidence organization, semantic inconsistency between visual and textual findings, and diminishing visual influence as generation progresses. We propose **VESP**, a visual evidence preservation framework for structured MRI report generation that maintains image-grounded evidence from visual encoding to final token prediction. First, **Schema-guided Visual Evidence Linking (SVEL)** organizes shared MRI representations into field-specific visual memories and propagates information among related clinical fields using schema-derived relations. Second, **Visual-Text Evidence Alignment (VTEA)** aligns image-side and report-side predictions within a shared structured finding space, encouraging generated findings to remain consistent with visual evidence. Third, **Evidence-Retained Adaptive Decoding (ERAD)** selectively reuses learned field-level visual predictions at schema-defined candidate decisions during inference, preserving visual influence without additional training or learnable parameters. Experiments on a private shoulder MRI dataset and the public KneeCoT knee MRI dataset show consistent improvements, achieving 75.65 micro-F1 on shoulder MRI and 96.69 micro-F1 on knee MRI. These results demonstrate that preserving evidence identity, semantics, and influence throughout autoregressive generation improves the accuracy and image consistency of structured MRI reports.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.