Learning Where to Look: Reinforcement Learning from Radiologist Gaze for Visually Grounded Chest X-ray VQA
Abstract
Recent medical vision-language models have shown strong capability in chest X-ray (CXR) interpretation. Supervised fine-tuning (SFT), which optimizes the answer tokens, has enabled this capability through instruction tuning, and reinforcement learning (RL) has further improved it by updating the policy with rewards on answer correctness or reasoning quality. However, these training signals supervise only the generated text rather than how the model reaches the answer from the image. Neither objective constrains where the model looks, allowing a model to answer correctly while attending to clinically irrelevant regions. We propose CheXGaze, a CXR-specialized VLM for visual question answering that learns where to look by using radiologists' eye gaze as an RL reward rather than a target to imitate. Following radiologists' global-to-focal reading process, CheXGaze generates an <organ> bounding box, <gaze> points, and an <answer>. Experiments on MIMIC-CXR-VQA and the CXR subsets of SLAKE and VQA-RAD demonstrate its effectiveness, and attention analysis shows closer alignment with radiologists' gaze, indicating that CheXGaze learns where to look.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.