acceptodds
Under review as a conference paper at ICLR 2027

TumorCast: A Cancer World Model for Longitudinal Tumor Forecasting

Abstract

Anticipating how a tumor will grow and spread could help clinicians adapt therapy in time, yet forecasting its evolution is difficult: tumors exhibit substantial heterogeneity, and their trajectories depend on anatomical location, current disease state, local tissue context, and the patient’s clinical history. Existing approaches often predict clinical outcomes directly from observed CT scans or records, or generate future CT scans, leaving the evolution of CT latent states insufficiently modeled. To address this gap, we propose TumorCast, a cancer world model for longitudinal tumor forecasting that predicts CT latent states conditioned on clinical history, interval treatment, and transition interval. Built on the joint-embedding predictive architecture (JEPA), TumorCast learns reusable predicted CT latent states through a Representation–Dynamics–Readout curriculum, without reconstructing future CT scans. First, longitudinal CT scan pairs adapt a CT encoder and latent predictor through future-state supervision, while clinical text provides conditioning inputs for state transitions. Next, chained training learns successive state transitions through recursive rollout, with future-state supervision propagating backward through the trajectory so that later prediction errors refine earlier state transitions and predicted intermediate states. Finally, a shared clinical reader (Qwen3.5-9B/27B) combines predicted CT latent states with clinical context to forecast lesion diameter, lesion size change, and liver metastasis status. To support evaluation, we curate a multimodal tumor forecasting benchmark comprising 1,344 patients, 3,606 CT scans with corresponding radiology reports, and 11,728 structured clinical entries covering clinical history and interval treatment, yielding 2,400 forecasting examples across liver and pancreas. Experiments show that TumorCast outperforms rendered-slice vision-language models (VLMs) and text-only baselines, highlighting the success of our multimodal JEPA design and potential clinical effectiveness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.