PAIR3D: Learning Physics-Aligned and Visually Consistent 3D Scene Layouts from Verifiable Rewards
Abstract
Modern 3D reconstruction can produce scenes that look right but behave wrong. Small errors in object layout or geometry may remain inconspicuous in the input view, yet become explicit under rigid-body dynamics through falling, toppling, interpenetration, or unstable contact responses. Recent methods use physical simulation to improve such reconstructions, but their scene-layout decisions are still largely worked out anew for each input through test-time search or prescribed diagnostic procedures. We ask whether this alignment behavior can instead be learned and reused. Our key observation is that the right action can be ambiguous even when its outcome is directly verifiable. We therefore introduce PAVER3D, which formulates physical scene alignment as sequential decision making with verifiable rewards. A multimodal relational policy reasons over current and target visual evidence, object states, and pairwise physical relations, and predicts objectlayout updates. Simulator outcomes train the policy with PPO, amortizing part of the per-scene optimization into reusable behavior while retaining a closed-loop simulator at deployment. In our current comparison with REST3D, PAVER3D improves physical stability more strongly, changes mIoU by +1.0% rather than −1.6398% relative to the initial reconstruction, and reduces measured alignment time from 42.23 to 4.50 seconds per scene. These results suggest that simulation can serve not only as a scene-specific optimizer, but also as a source of experience for learning how to align 3D scenes with physics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.