FAME: Factorized Astronomical Multimodal Embeddings
Abstract
Stellar astrophysics observes the same stars through different instruments, across surveys with vastly different noise, resolution, and coverage, and some key properties, like stellar age, are not measured by any single one of them. Building a foundation model that combines these modalities therefore remains an open challenge. We present FAME, a foundation model for stellar astrophysics that combines seven modalities from six large-scale surveys via pretrained modality-specific encoders and a small alignment module. FAME departs from classical multimodal alignment in two ways. Its latent space is factorized into a shared and a residual subspace, so that aligning the modalities preserves what each one holds alone. And the shared subspace is trained with a joint-embedding objective extended by a graph-Laplacian penalty over catalogued physical coordinates. Under frozen linear probing, FAME outperforms unaligned features, a single-latent alignment baseline, and three existing astrophysics foundation models. Against fully supervised models it is more label-efficient on the labels that enter no pretraining objective and is beaten only on the one label read directly from a single observational capability; when the surveys carrying such a label are withheld, FAME still recovers it ( and ), where a supervised model trained under the same schedule retains little ( and ). Rare stellar populations are linearly decodable from the residual subspace, and the representation supports cross-modal spectral generation. Neither ingredient is specific to astronomy: the factorization applies wherever modalities carry private information, and the smoothness penalty wherever samples come with catalogued continuous coordinates. To our knowledge, FAME is the first foundation model for stellar astrophysics to combine time series, spectra and astrometry at scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.