acceptodds
Under review as a conference paper at ICLR 2027

Learning Where to Look: Reinforcement Learning from Radiologist Gaze for Visually Grounded Chest X-ray VQA

Abstract

Recent medical vision-language models have shown strong capability in chest X-ray (CXR) interpretation. Supervised fine-tuning (SFT), which optimizes the answer tokens, has enabled this capability through instruction tuning, and reinforcement learning (RL) has further improved it by updating the policy with rewards on answer correctness or reasoning quality. However, these training signals supervise only the generated text rather than how the model reaches the answer from the image. Neither objective constrains where the model looks, allowing a model to answer correctly while attending to clinically irrelevant regions. We propose CheXGaze, a CXR-specialized VLM for visual question answering that learns where to look by using radiologists' eye gaze as an RL reward rather than a target to imitate. Following radiologists' global-to-focal reading process, CheXGaze generates an <organ> bounding box, <gaze> points, and an <answer>. Experiments on MIMIC-CXR-VQA and the CXR subsets of SLAKE and VQA-RAD demonstrate its effectiveness, and attention analysis shows closer alignment with radiologists' gaze, indicating that CheXGaze learns where to look.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.