acceptodds
Under review as a conference paper at ICLR 2027

CoRe3D: Collaborative Reasoning through Octant-Based Chain-of-Thought for Unified 3D Generation and Understanding

Abstract

Recent progress in multimodal foundation models shows that explicit reasoning mechanisms can improve reliability, cross-modal alignment, and generation, yet comparable mechanisms for 3D remain limited. We present CoRe3D, a unified framework for 3D understanding and generation, that performs collaborative reasoning over two coupled traces: (i) a semantic CoT that expands a prompt into an interpretable plan (parts, attributes, relations), and (ii) an octant-based geometric CoT that realizes this plan as localized 3D tokens, enabling compositional and interpretable generation. To jointly optimize both traces end-to-end under heterogeneous, partly non-differentiable 3D objectives, we propose Co-GRPO, a trace-aware group-relative preference optimization that integrates multi-critic 3D rewards. Experiments show that CoRe3D improves 3D understanding (e.g. +27% BLEU-1, +23% ROUGE-L, +13% METEOR over the strongest baseline) and achieves state-of-the-art text-to-3D alignment (+38% CLIP over the next best method), while also supporting a diverse set of tasks spanning general multimodal reasoning, image-to-3D, reasoning-based 3D generation over complex prompts, and fine-grained 3D part editing within a single model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.