A Structured Commitment Space for Factual Policy Learning in Chest Radiography
Abstract
Clinical vision-language models can identify the correct finding yet misstate whether it is present or how severe it is. Aggregate reinforcement-learning rewards provide little explicit control over how these different factual errors are penalized. We introduce Fact-Commitment Space, a reward for factual policy learning in chest radiography. It represents free-form clinical text as commitments, each specifying a clinical entity, its assertion state, and relevant attributes. A Clinical Fact Parser (CFP) extracts these commitments, and a deterministic compiler normalizes them. Matching response and reference commitments then assigns separate penalties to omissions, unsupported findings, attribute errors, and contradictions, without an online language-model judge. Under stated conditions on structured commitments, omitting an attribute has lower expected loss than guessing in the zero-confidence limit, reversing an assertion scores below omission, and adding an unsupported commitment lowers the reward. On 600 human-reviewed responses, CFP improves entity–assertion F1 from 13.82% to 75.32% over its Qwen3.5-2B base, and its rewards preserve 72.78% of non-tied reference pair orderings. Training with this reward using Group Relative Policy Optimization (GRPO) yields VeraRad-2B. On 3,000 clinical visual questions, it achieves 75.53% strict accuracy versus 66.70% for its Qwen3.5-2B base, with thinking enabled for both models. On report generation, a task absent from training, it leads the primary quality metric on two of three datasets and produces reports 43.7% shorter than its base in the same mode. Image-swap and no-image interventions assess dependence on study-specific images, complemented by spatial attribution analysis of a cardiomegaly case.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.