Think Again and Look Closer: Dual-Exploration Test-Time Reinforcement Learning for Medical Multimodal Reasoning
Abstract
Medical Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across healthcare tasks, yet their post-training strategies still heavily depend on substantial annotated data, overlooking the potential of large-scale unlabeled clinical data. An alternative approach, test-time reinforcement learning, operates by sampling multiple trajectories for each test input and exploiting the consensus among their answers as a self-supervised signal. However, for medical visual question answering (VQA), consensus among reasoning answers alone does not constitute reliable evidence. From a textual perspective, over-reliance on answer consensus can amplify confidence in unsubstantiated conclusions and may even encourage shorter, less informative responses that shortcut reasoning in favor of textual agreement. Visually, when multiple sampled trajectories converge on the same answer, their consensus should be grounded in consistent visual evidence rather than textual consistency alone. Clinically, physicians focus intently on subtle abnormalities, whereas even with attention mechanisms, general-purpose models are distracted by large irrelevant areas in the whole image. These limitations call for test-time learning improvements that promote consensus without textual shortcuts, visual grounding consistency across trajectories, and localized exploration beyond holistic features. We therefore introduce TALC, a dual-exploration test-time reinforcement learning framework that asks the model to Think Again and Look Closer. Think Again first anchors on answer support, then ranks reasoning paths by semantic novelty within each support group while also measuring visual consistency across trajectories. Look Closer, in turn, through explicit and executable region-selection actions, assigns each region proposal a tool-token uncertainty score, and branches from that decision point to resample alternative local inspections. Experiments across multiple medical VQA benchmarks and MLLM backbones demonstrate the effectiveness and generalizability of TALC for test-time medical multimodal reasoning. Our code is available at: https://anonymous.4open.science/r/TALC.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.