HIdden in the DEcodes: Recovering Clinical Findings without Retraining Medical MLLMs
Abstract
Automatic radiology report generation commonly commits to a single decoded report from a medical multimodal large language model (MLLM). However, clinically correct findings omitted from this report may be hidden in other stochastic decodes produced by the same frozen model, while unsupported findings may also recur across decodes. Existing approaches address these errors by modifying training objectives, model architectures, supervision, or refinement modules, but typically require model updates or additional training and do not ensure that a single decode captures all supported findings. We introduce HIDE, a model-agnostic, training-free framework for recovering clinical findings without retraining medical MLLMs. HIDE treats the frozen generator as a proposal distribution for verification-guided inference: stochastic decoding broadens candidate coverage, cross-sample consistency provides model-side proposal support, localized visual support from a prototype bank verifies image compatibility, and source-constrained minimal edits compose only jointly supported findings. Given an initial draft and additional stochastic samples, HIDE extracts finding-level candidates and selects among three conservative actions: adding a supported finding missing from the draft, removing an unsupported draft claim, or abstaining when the evidence is inconclusive. Across six frozen MLLM report generators, HIDE achieves mean relative gains of 27.3%, 54.1%, and 16.0% on CheXbert-F1-14, CheXbert-F1-5, and RadGraph-F1 on MIMIC-CXR, respectively. On CheXpert Plus, the corresponding gains reach 41.2%, 239.7%, and 19.4%. The correction signal also transfers to IU-Xray when fitted only on MIMIC-CXR or CheXpert Plus. HIDE further achieves state-of-the-art performance among training-free correction methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.