SIA-CT:State-Structured Item Alignment for Open-Vocabulary 3D Chest CT Representation Learning
Abstract
Radiology reports describe both findings and the states that determine their meaning: where a finding occurs, whether it is present, and how severe or certain it is. Learning from these reports requires supervision that preserves these distinctions. We introduce SIA-CT, a framework for state-structured item alignment in 3D chest CT. Its central principle is that cross-patient item alignment requires agreement in observed states, whereas within-examination item separation requires evidence of distinct findings. We instantiate this principle through coordinated item separation, cross-patient alignment, and false-negative masking. A complementary regional objective aligns organ-supported visual features with report descriptions through shared transformations. Together, these objectives train a representation queried by natural language using only CT and text at inference. On 750 report-derived state questions, SIA-CT reaches 66.8% four-axis macro accuracy, exceeding a matched itemized baseline by 3.5 percentage points. Global and textconditioned score fusion yields AUROCs of 83.3% on CT-RATE and 75.3% on RadChestCT, while text-conditioned scoring gives stronger aggregate state discrimination. These results establish observed-state relations as a useful source of supervision and show why finding recognition and state discrimination should be evaluated together.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.