RAPP3R: Reinforcing Object Relationship for Physically Plausible 3D Reconstruction
Abstract
We present RAPP3R, a unified framework that combines architectural enhancements with a reinforcement learning strategy for physically plausible multi-object 3D scene reconstruction from a single image. Building on a pre-trained joint shape-pose generator that processes each object independently, RAPP3R equips the model to reason about interactions across the objects. Gated Inter-Object Attention introduces an inter-object branch while preserving the pre-trained model's behavior at initialization. Volume-RoPE further encodes each object's position and spatial extent, providing explicit spatial context for attention across objects. Building on these architectural changes, we train the model with GDPO to learn physical relationships in a scene, such as collisions and floating. We overcome the limitations of scene-level reinforcement learning through object-wise advantage computation. As a result, our experiments show that RAPP3R outperforms prior methods in both pose accuracy and physical plausibility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.