acceptodds
Under review as a conference paper at ICLR 2027

DynaRep: Learning Future Visual Representations for Object-Centric Dynamics

Abstract

Most embodied world models predict futures in latent spaces optimized for pixel reconstruction, which may not be ideal for capturing manipulation-relevant dynamics. We present DynaRep, a world model that represents future world states in a visual representation space rich in semantic and geometric structure. To this end, we jointly model future DINO representations and 3D object flow to capture object-centric dynamics. To make dense DINO features tractable for generative modeling, we introduce DINO-VAE, which performs channel and temporal compression with causal temporal modeling to produce compact latents while preserving informative visual structure. Building on this compact representation space, we employ a unified diffusion transformer with joint attention to capture the bidirectional dependencies between future visual representations and 3D object flow. To further strengthen cross-modal coupling, DynaRep adopts asymmetric modality-specific noise schedules, allowing visual representations to stabilize earlier during denoising and continuously guide 3D object flow generation. Experiments on simulated and real-world manipulation data show that DINO-derived representations provide a more effective space for future modeling than pixel-reconstruction-oriented latents, leading to more geometrically consistent and task-relevant 3D object flow. These results demonstrate the value of predictive visual representations for modeling object-centric dynamics in embodied interaction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.