FACET: Rethinking Whole-Slide Foundation Models Through Task-Conditioned Views
Abstract
Foundation models (FMs) have demonstrated great success in computational pathology, first at the patch level and more recently at the whole-slide level. Despite being trained on increasingly large WSI corpora and, in some cases, incorporating additional molecular or textual modalities, existing slide-level FMs still produce a single static representation per slide, agnostic to the downstream task. This design is inherently limiting: a whole-slide image contains diverse morphological signals, only a subset of which may be relevant for any given clinical or biological question. We propose FACET, a whole-slide foundation model that conditions slide representations on task descriptions, enabling dynamic inference-time aggregation tailored to each prediction objective. FACET is pretrained through supervised multitask learning across 29 carefully curated prediction tasks spanning six cancer types, aligning task-conditioned slide representations with text-derived class prototypes. We evaluate FACET on an external benchmark comprising 40 distinct tasks, including tasks and cancer types not observed during pretraining. FACET achieves state-of-the-art performance against leading slide-level foundation models and maintains high performance on both unseen tasks and unseen cancer types, demonstrating robust generalization. Together, these results demonstrate task conditioning as a powerful approach for learning generalizable whole-slide representations. We release code, label files, task manifests, and evaluation resources at https://github.com/anonymous-95/FACET.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.