acceptodds
Under review as a conference paper at ICLR 2027

DreamBridge: Bridged Cross-Modal Distillation for Text-to-3D Generation

Abstract

Existing text-to-3D generation models can produce semantically consistent 3D ob- jects from natural-language prompts, yet they still struggle to capture the geometry and spatial layout of complex objects. Common approaches either train native 3D generators on large paired datasets or combine text-to-image synthesis with image- to-3D generation at inference time. Inspired by the latter, we propose DreamBridge, a bridged cross-modal distillation framework that transfers sparse-structure priors from a frozen image-to-3D teacher to a text-to-3D student, where synthesized im- ages bridge the conditioning modalities during training. DreamBridge comprises three components: (i) Bridged Triplets Synthesis, which constructs triplets of a prompt, a filtered FLUX-generated image, and a cached teacher sparse-structure la- tent; (ii) Native Latent Distillation, which adapts only the student’s sparse-structure flow with LoRA through conditional flow matching on cached teacher latents; and (iii) Anchor Replay, which supplements teacher supervision with structural samples from the frozen original text model. The overall adaptation procedure uses model-generated supervision without additional ground-truth 3D training data and retains text-only inference. On BridgeStruct-1.5K, our evaluation suite of 1,536 held-out prompts, DreamBridge improves structural agreement with the teacher over the pretrained baseline and achieves the highest reported text–3D semantic alignment scores among the evaluated native text-to-3D models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.