acceptodds
Under review as a conference paper at ICLR 2027

ImaGeo: Bidirectional Image–Geometry Co-Reasoning for 3D Generation

Abstract

Image-conditioned 3D generation requires fidelity to a single image’s visible structure and details, and consistency with complementary observations when multiple views are available. We present ImaGeo, a diffusion transformer that generates 3D geometry from one or more input images. Its denoising blocks combine per-view image and geometry self-attention with joint attention, connecting all views to a shared geometry representation. Updated image features carry geometric and cross-view context into subsequent blocks while retaining per-view spatial processing. The same design operates in coarse and fine generation, trained solely through geometry flow matching without a dedicated reconstruction branch or auxiliary reconstruction losses. It supports variable-view inputs without input camera parameters. On GSO and Toys4K, ImaGeo improves single-view image–geometry consistency and sparse-view reconstruction accuracy. Evaluations on AI-generated images also show improved image–shape similarity. In controlled single-view fine-stage comparisons, our 1.5B denoiser outperforms a parameter-matched cross-attention baseline and achieves normal-map quality comparable to a 5B baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.