acceptodds
Under review as a conference paper at ICLR 2027

SceneProbe: Reconstructing and Physically Testing Scenes from Images and Videos

Abstract

A scene reconstructed from an image or video must withstand gravity and controlled actions, yet a plausible render cannot show whether its objects move independently, a hinge is present, a cushion has a filled volume, or transient collisions occur. We present SceneProbe, an image/video-to-MuJoCo pipeline that keeps each source instance linked to a native physical component and tests the exact exported model. Its Instance Evidence Graph records masks, geometry, placement and support/part cues; Type-Aware Native Assembly creates separate rigid bodies, fitted joints and, when feasible, contact-enabled tetrahedral objects with declared parameter priors; and Intervention-Aware Native Audit checks realized structure and every integration state under gravity and forces. In a selected corrected living-room image scene, 13 predicted targets include two functional hinges and two volumetric pillows, and the five-second rollout reaches 1.825 mm peak enabled-contact penetration. Corrected toy tabletop image and video scenes also run, although the video route remains mainly first-frame conditioned. Across 68 dependent development XMLs, saved 10 Hz states miss 14 of 38 every-step contact violations at a 2 mm operational threshold. None of ten frozen public-image first attempts passes the combined coverage, completion and contact gate; four authored source-truth RGB scenes also fail complete structure/action recovery. SceneProbe demonstrates auditable native reconstructions in selected cases while exposing an automation gap. Mass, friction and elasticity are priors, not recovered real-world measurements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.