Holi-World: Building Holistic World Intelligence from Human-Ego Experience
Abstract
Scaling spatial intelligence requires object-centric supervision that remains valid not only across viewpoints, but also as the physical world changes over time. Existing 3D data pipelines largely assume static scenes: they either depend on costly human-annotated scans or collapse a video into a single scene, discarding object motion. Conversely, applying a 3D detector independently to each frame produces noisy boxes, fragmented identities, and inconsistent trajectories. We present SPAFoundry, a unified and scalable pipeline for open-vocabulary static and dynamic 3D grounding from posed visual sequences. SPAFoundry uses Boxer as a shared 2D-to-3D proposal generator and introduces complementary scene-level consolidation paths. The static path aggregates repeated observations in a common world frame to refine object geometry and suppress duplicates, while the dynamic path preserves time, associates moving instances across frames, and recovers temporally consistent 3D boxes and trajectories. The resulting interface supports heterogeneous image sequences with optional sparse or dense geometric cues and exports a common annotation format for static objects and dynamic tracks. On ScanNet, SPAFoundry improves AP from 0.32 to 0.68 over Boxer. On 20 Aria Digital Twin sequences, it achieves 30.88% dynamic 3D box IoU, 54.31% HOTA, and a TP-L2 error of 0.0333 m. These results show that explicitly modeling static persistence and dynamic motion turns robust 2D-to-3D lifting into an effective data engine for multimodal spatial intelligence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.