What Makes a Latent Space Shared? Direct and Recoverable Compatibility in Motion VAEs
Abstract
A shared latent space is often assumed to make representations from different schemas directly interchangeable, yet sharing parameters, dimensionality, or a prior does not determine how corresponding instances are expressed in latent coordinates. We show that this assumption conflates recoverable correspondence, direct compatibility, and decoder reuse. In a controlled multi-schema VAE on HumanML3D, even identical inputs yield only – raw Recall@1 and – paired-token cosine, despite correspondence being almost perfectly recoverable after linear calibration. Cycle consistency strengthens recoverability but still leaves direct retrieval and token agreement near zero. In contrast, only paired Rotations clips raise raw Recall@1 to – and token cosine to approximately , making most post-hoc alignment unnecessary, although crossed decoding remains substantially worse than native reconstruction. We further find that reducing measured marginal discrepancy does not imply schema invariance. These results show that a shared backbone or prior does not ensure directly compatible latent coordinates, while limited paired correspondence can substantially improve direct compatibility in the tested Rotations setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.