Seeing Is Not Saying: Closing the Verbalization Gap in 3D CT Report Generation
Abstract
Recent work on 3D Computed Tomography (CT) radiology report generation has largely focused on improving visual representations and the evidence supplied to language models. We identify a substantial gap between what is recoverable from visual representations and what is ultimately expressed in generated reports. On the example of thorax abnormalities, lightweight probes across five frozen encoders recover substantially stronger abnormality signals than report generators express, including for encoders not trained with CT reports. We term this discrepancy the verbalization gap and address it by improving both the evidence available for generation and how that evidence is expressed. Our framework consolidates complementary CT representations by slot-wise convex weights learned over independently resampled tokens and introduces graded verbalization, which guides report generation with diagnostic predictions whose linguistic certainty reflects their class-level reliability. With a 3B language backbone, our system achieves 0.597 macro F1 and 0.526 CRG on CT-RATE, a relative improvement of 44% over AdaRAG-CT and 180% over BTB3D. Under zero-shot cross-dataset evaluation on RAD-ChestCT, it achieves 0.498 macro F1, 123% above BTB3D. These results indicate that 3D CT report generation is limited not only by what can be recovered from visual representations, but also by how effectively recoverable clinical findings are translated into language.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.