MedJEPA: Latent Cross-Modal Prediction for Medical Vision-Language Reasoning under Missing Images
Abstract
Vision Language Models (VLMs) have become effective tools for reasoning over medical images, yet their multimodal reasoning pathways typically assume that the image is available at inference. In many real-world clinical scenarios, images may be unavailable because they are not transmitted across clinical workflows, are withheld due to privacy or data-sharing constraints, are corrupted during transmission, or have not yet been acquired. Existing missing-modality methods rely on prompt adaptation, retrieval, or modality reconstruction, but these strategies either work around the missing input or attempt to recreate the observation itself, rather than directly recovering the latent visual representation that the downstream VLM was trained to reason over. To overcome these limitations, we propose MedJEPA, a latent cross-modal prediction framework inspired by the Joint-Embedding Predictive Architecture (JEPA). Given only the clinical text and the corresponding question, MedJEPA predicts the patch-level visual representations that a frozen medical VLM would extract from the missing image, without synthesizing pixels.The predictor is trained with a JEPA-style masked latent objective in CLIP feature space, using the frozen image representation as a stop-gradient target. A masking curriculum spans partial to complete visual absence, with patch annealing progressively removing real visual tokens and iterative self-refinement improving the latent prediction from previous estimates. At inference time, the predicted visual tokens are projected through the frozen multimodal projector of the VLM and injected into its visual stream. A confidence-gated cascade then progressively evaluates sparse and dense visual-token injection, reverting to text-only reasoning when the predicted visual signal fails to improve the model’s confidence. Across PMC-VQA and ScienceQA, MedJEPA consistently improves image-absent accuracy, from 70.8% to 73.4% on PMC-VQA and from 68.8% to 72.3% on ScienceQA, yielding gains of 2.6 and 3.5 percentage points, respectively. Ablations confirm the benefits of latent prediction, anti-collapse regularization, and self-refinement. These results show that latent cross-modal prediction enables medical VLMs to reason more reliably when images are unavailable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.