HOME: Multi-Room Geometry Estimation via Top-Down Guidance
Abstract
Feed-forward 3D foundation models struggle in multi-room environments where cameras are segregated across rooms divided by opaque walls, breaking visual overlap and spatial continuity. We propose HOME, a method that adapts 3D foundation models to incorporate an unposed top-down RGB render alongside perspective captures for multi-room geometry and relative camera pose estimation. To support this, we construct TD-HOMES, a dataset suite spanning real-world scans, designer-crafted spaces, and agentic procedural environments. While providing global top-down context largely resolves catastrophic pose and geometry failures across disjoint rooms, naively sharing parameters induces representational interference that compromises pre-trained metric precision on standard overlapping scenes. To alleviate this issue, HOME introduces a view-type decoupling design that selectively decouples the projection weights and feed-forward networks within global cross-frame attention layers between top-down and perspective views. Experiments show that HOME reliably recovers coherent multi-room geometry across non-overlapping views while reducing performance drops on standard overlapping scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.