SceneWise: Learning a Decision-Consequence World Model for 3D Indoor Scene Generation
Abstract
Recent agentic frameworks for 3D indoor scene generation integrate generation, placement, and editing tools to construct scenes from natural language instructions, refining intermediate layouts through feedback. However, they overlook how currently valid decisions can prevent subsequent spatial and functional requirements from being met, leading to delayed failures and costly revisions. We propose SceneWise, a framework that learns a decision-consequence world model to guide scene generation and repair. During generation, the world model predicts how candidate decisions affect placement feasibility and constraint satisfaction, guiding agents to avoid downstream conflicts. During repair, an attribution model uses the failure context and historical decision graph to rank earlier decisions for revision. The same world model evaluates replacement candidates to guide local regeneration while limiting changes to valid scene content. We further introduce SceneWiseDB, a dataset of 100k records across 20 room categories that provides consequence and intervention supervision for both models. Experiments show that SceneWise improves task success and constraint satisfaction while reducing repeated editing, and enables effective repair with limited changes to existing scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.