TACER: Verified 3D Tabletop Scene Generation for Simulation and Interaction
Abstract
Generating 3D tabletop scenes from images or text supports robot learning, but visual and geometric plausibility alone do not establish simulation validity or interaction usability. Existing methods largely operate on static meshs, with limited support for functional structures. Evaluation also emphasizes static geometry and visual judgments, offering limited evidence of dynamic validity and task executability. To address these limitations, we present TACER, which connects asset construction, scene assembly, and verification-guided refinement through hierarchical executable programs. Asset-level programs construct geometry and physical representations, including articulation; scene-level programs organize placements and spatial relations. Geometric checks and simulation tests guide asset or layout revisions, which are re-executed and verified before acceptance. To independently assess scene capabilities, we introduce TabletopAudit, a 500-case benchmark covering input fidelity, 3D scene quality, simulation validity, and robot task execution. TACER leads input fidelity on both input tracks. On Embodied-100, it achieves 71.67 rigid-manipulation and 60.00% articulated task success, compared with 20.00% for TabletopGen and 5.00% for TabletopGen + Articraft, respectively. To demonstrate scalability, we construct TACER-11K: 11,399 scenes across nine environments with physical representations, interactive assets, and available verification evidence. A downstream study demonstrates their utility for demonstration collection and task-specific policy learning. Together, these results support TACER's effectiveness and establish TabletopAudit as a basis for assessing simulation readiness and interaction usability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.