ROSE-3D: Render-Free Object-to-Scene Embeddings for Compositional 3D Representations
Abstract
3D scene representations are typically computed from the entire scene at a global level, rather than from reusable representations of the objects it contains. We argue that this constrains their capacity to express composition. A scene consists of objects placed in a shared frame, and the same objects recur across scenes; encoding each scene as a whole precludes reuse, making new arrangements of familiar objects as costly as unfamiliar ones. We treat composition as an operation in representation space: a learned operator maps object representations and their placements (translation, rotation, and scale) to a representation of the scene, without assembly or global encoding. To keep an object's content (shape and appearance) separable from its placement, we encode each 3D asset from its mesh and material maps without rendering, binding local surface orientation and appearance to continuous surface coordinates. We show that predicting a scene encoder's features can be satisfied without reading object content: the operator relies largely on layout and barely changes when objects are substituted. We therefore add two complementary objectives: identification, which requires the representation at each location to identify the object placed there, and language supervision, which aligns the composed representation with a description of the scene. With these objectives, the composed representation follows object content: it responds to substituted objects, including objects unseen during composer training, and after a replacement it shifts toward the new object. On question answering and situated next-step navigation, composed representations are competitive with point-map and point-cloud methods and outperform simpler representations built from the same inputs, while representing each new arrangement in milliseconds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.