Post-Training Multimodal Embedding Models towards Physical Dynamics Awareness
Abstract
Physical AI needs to connect what is observed with what can happen next, but pixel-level prediction couples learning physical dynamics with reproducing potentially irrelevant appearance details. Predicting learned representations can ease this burden while retaining the state information needed to anticipate change. Embedding models built on multimodal large language models (MLLMs) offer a promising starting point, with our experiments indicating greater utility as physical state encoders than same-size MLLMs. These embedding models combine vision–language processing with reusable representations, yet their relevance-oriented training does not directly target physical transitions. We use soft tokens and learned-query pooling to learn physical state representations alongside a frozen backbone's native retrieval embeddings, then look beyond task-specific adaptation to explore cross-task predictive post-training of a shared, dynamics-aware encoder. Frozen-backbone adaptation delivers stronger overall planning performance than LeWorldModel with fewer trainable parameters, while cross-task post-training improves the shared encoder's utility for world modeling, supporting the potential of a dynamics-aware representation foundation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.