SIGMA: Spatiotemporal Interventional Generation via Mechanism-Aware Unified Models
Abstract
Clinical management of blinding fundus diseases is a long-term process involving continuous follow-up and dynamic intervention. Beyond assessing the current pathological status, clinicians also need to evaluate how lesions may evolve over a future time horizon under a given treatment strategy. Existing medical vision AI is mainly organized around two task formulations: static perception, which focuses on the current pathological state in single-time-point images, and \medical image editing, which focuses on controlled image modification under given target attributes or editing constraints. These two paradigms address current-state analysis and target-appearance control, respectively, but neither explicitly models the relationship among treatment strategy, time horizon, and pathological evolution, making them unsuitable for future image forecasting under intervention constraints. To address this gap, we introduce Spatiotemporal Interventional Generation (SIG) Task. Given a baseline image, intervention variables, and a time horizon, SIG requires the model to jointly predict the future image and its structured clinical rationale. For this task, we propose SIGMA, a tool-augmented unified generation-understanding model. SIGMA first extracts discrete biomarker representations from the baseline image and uses them to filter an external memory, where candidates are ranked by image similarity. The retrieved trajectories, intervention, and time horizon then condition unified forecasting. Under such priors, the unified model jointly generates the clinical rationale and the future image. Experiments reveal the challenges that SIG poses to existing models and demonstrate the effectiveness of SIGMA after task-specific fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.