Look Beyond the Question: Preserve–Focus–Fuse for Whole-Slide Image Question Answering
Abstract
A question about a whole-slide image rarely names every morphological finding needed to answer it. Question-guided patch selection can therefore overlook supporting morphology, while slide-level compression can dilute focal evidence. Inspired by how pathologists combine broad slide review, question-directed examination, and integration of their findings, we introduce BeyondPath, a Preserve–Focus–Fuse framework for whole-slide image question answering. Preserve learns a question-agnostic Morphology Bank through answer supervision across multiple questions about each training slide. Focus retrieves question-specific evidence from all foreground patches, including regions outside the Bank. Fuse integrates the preserved morphology at the retrieved evidence positions through Morphology-on-Visual-Evidence (MoV) Fusion, producing a fixed 96-token visual prefix. On WSI-Bench, BeyondPath achieves an Avg-9 of 0.769 and leads all eight reported report-generation metrics using the same checkpoint. It also achieves 0.639 zero-shot overall accuracy on SlideBench-BCNB. Controlled ablations show complementary benefits from morphology preservation and question-directed retrieval, and support cross-question supervision and fusion at the evidence positions. These results support looking beyond the question’s explicit cues: preserving broader slide morphology while retaining the ability to seek evidence for each question.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.