acceptodds
Under review as a conference paper at ICLR 2027

PATHOARGUS: ADVANCING COMPREHENSIVE, EVIDENCE-GROUNDED VISUAL REASONING FROM GIGAPIXEL SLIDES TO MULTI-SLIDE CASES

Abstract

Whole-slide pathology question answering requires finding relevant tissue and an- swering from visual evidence, yet textual cues can also yield correct answers. We introduce PATHOARGUS-BENCH to evaluate six pathology capabilities and test whether answers follow controlled changes in target evidence. It contains 22,078 four-choice questions covering 4,809 target cases across 15 TCGA projects, in- cluding 4,281 bench questions. Evidence-Set Grounding (ESG) forms counterfac- tual quartets by placing a target in each of three candidate sets or removing it; quar- tet exact-match accuracy (QExact) requires all four answers to be correct. Across 24 baselines, GPT-5.6 attains 57.09% overall accuracy but only 3.93% QExact on 483 quartets. We further propose PATHOARGUS, a framework that selects question-relevant evidence patches from whole-slide images to answer pathology questions. It reaches 65.15% overall accuracy, 68.27% ESG accuracy, and 43.27% QExact using only 512 selected patches per question, retaining 1.51% of avail- able bench features. These results show more consistent answers across evidence states, while fine-grained local findings remain challenging.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.