Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
Abstract
Recent multimodal large language models have achieved remarkable success in unified text and image understanding and generation, yet extending such native capability to 3D remains an open challenge. The core bottleneck is data: compared to the near-infinite 2D imagery on the web, high-quality 3D assets are orders of magnitude scarcer, leaving 3D synthesis severely under-constrained. Current methods often circumvent this limitation through indirect pipelines that edit in 2D image space and lift results into 3D via iterative optimization, sacrificing geometric consistency and the straightforwardness of native generation. We present Omni123, a 3D native foundation model that addresses limited 3D data by unifying text-to-2D and text-to-3D generation within a single autoregressive framework. Our key insight is that cross-modal generative consistency between images and 3D can act as an implicit structural constraint: by representing text, images, and 3D geometry as discrete tokens in a shared sequence space, the model can leverage abundant 2D observations as a rich geometric prior to fortify 3D representations. To realize this, we introduce an interleaved X-to-X training paradigm that coordinates diverse cross-modal tasks over heterogeneous paired datasets without requiring fully aligned text–image–3D triplets. By traversing semantic–visual–geometric cycles (e.g., text → image → 3D → image) within single autoregressive sequences, the model learns representations that simultaneously satisfy high-level semantic intent, appearance fidelity, and multi-view geometric consistency, while mitigating harmful interference between appearance and geometry objectives. Extensive experiments demonstrate that Omni123 significantly improves geometric consistency and semantic alignment in text-guided 3D generation and editing, validating that unifying 2D and 3D generative processes provides an effective and scalable pathway toward multimodal 3D world models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.