acceptodds
Under review as a conference paper at ICLR 2027

Learning Temporally Equivariant Multi-Modal Large Language Models for Longitudinal Medical Image Comparison

Abstract

Longitudinal medical image comparison requires models to reason about how findings change over time. A reliable model should therefore behave consistently under temporal reversal: if a finding worsens from a historical study to a current study, reversing the image order should yield the corresponding inverse relation. We refer to this property as temporal equivariance. A controlled study on CheXTemporal shows that current general and medical multi-modal large language models (MLLMs) often violate this property: similar forward and reverse accuracies do not imply reliable bidirectional correctness. We propose TEMA, a Temporally Equivariant MLLM Adaptation framework for longitudinal medical image comparison. TEMA couples supervision on original and temporally reversed image pairs with a visual evidence consistency objective: temporal semantics should transform under reversal, while their supporting visual evidence should remain consistent. To account for different output structures, TEMA operates at different weighting granularities: sample-level weighting for closed-ended temporal VQA and temporal-token weighting for open-ended report generation. We evaluate our TEMA on temporal VQA and open-ended difference report generation across multiple MLLM backbones. On temporal VQA, TEMA consistently improves both balanced accuracy and strict temporal equivariance over standard LoRA fine-tuning and temporal-inversion baselines across all four evaluated backbones. On difference report generation, TEMA also improves clinically oriented generation metrics across backbones and remains competitive with specialized longitudinal report-generation baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.