SynOphsis: Unified Ophthalmic Image and Volume Generation Across Modalities and Dimensions
Abstract
An ophthalmic examination records several 2D modalities, a volumetric OCT scan at one of two slab thicknesses, the en-face image registered to it, and device and free-text clinical metadata. Assessment rests on more than one of them. Generative models of this evidence are built one modality or one dimensionality at a time, so covering an examination takes a collection of models with no shared representation, and not even the volume and the en-face image acquired with it are generated together. We present SynOphsis, a unified generative model for ophthalmic imaging. A frozen autoencoder maps every input into one latent space and per-shape patch embedders place images and volumes on a common token grid, so a volume and the en-face view acquired with it occupy one self-attention sequence, which a single stochastic-interpolant transformer denoises. Conditioning is a mask over that sequence, so any subset of the evidence in either direction is the same operation on the same weights, and one set of weights covers eleven 2D modalities, all three volume protocols, sparse-observation completion, and generation conditioned on a paired image, modality and device, with text conditioning shown qualitatively. The pair is sampled along a single trajectory and either half may be generated from the other, a mode no prior ophthalmic generator provides. One checkpoint serves every mode we evaluate, with no task-specific head and no fine-tuning, while each conditioning baseline carries a mechanism built for one task family and several modes are outputs those architectures cannot emit. On the tasks those specialists do support, the broader interface keeps their volumetric reconstruction accuracy, and joint tokenization costs under 0.01% of the parameters. Separately, our SiT-based 3D stack improves over the retrained volumetric latent-diffusion stack by 3.28 dB PSNR under matched observation conditioning, and on held-out OCT volumes SynOphsis-XL takes FVD from that baseline's 56.2 to 32.9, against a floor of 29.5 measured between two disjoint halves of the real data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.