acceptodds
Under review as a conference paper at ICLR 2027

WholeSight: Learning to Think with Whole Slide Images

Abstract

Reasoning over a whole slide image (WSI) requires collecting evidence from sparse tissue regions across spatial locations and magnifications. A WSI agent must decide where to look, what to inspect at higher magnification, and how new observations should guide subsequent reasoning. Existing WSI multimodal systems have improved either slide representations or active inspection, but visual evidence acquisition is typically fixed, prescribed at inference time, or learned separately from reasoning. We introduce WholeSight, which instead treats the WSI as an interactive environment in which visual search and reasoning are learned jointly as a single policy. A WSI harness exposes regional search and local inspection as executable actions, allowing the agent to acquire and interpret evidence over successive observations. Supervised fine-tuning provides a cold start for reasoning and tool use from complete interaction trajectories, after which reinforcement learning optimizes evidence acquisition and reasoning through the agent's own slide interactions. Across 2,956 questions spanning ten evaluation settings, WholeSight ranks first in eight and outperforms the strongest evaluated WSI system of comparable scale by 11.10 percentage points overall. Relative to its original multimodal large language model backbone under the same visual interface, the learned policy improves performance across every evaluation, including slide sources outside the training collection drawn from The Cancer Genome Atlas. Our code is available at https://anonymous.4open.science/r/anonymous-method-release-A53C/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.