FORGE3D: Faithful Observation-Grounded 3D Mesh Generation via Latent Forcing
Abstract
Recent 3D mesh generators produce visually compelling assets, yet faithfully reconstructing specific objects from real-world observations remains challenging. A practical mesh generation framework must jointly preserve observed 3D geometry and pixel-aligned appearance from one or more views while plausibly completing unseen regions. However, these observations often serve only as coarse conditioning signals, allowing generated geometry and appearance to deviate from the input. Moreover, existing 3D mesh benchmarks provide limited coverage of realistic scenes and challenging observation conditions needed to assess reconstruction faithfulness in real-world scenarios. To address these gaps, we introduce FORGE3D (Faithful Observation-Grounded 3D Mesh Generation), an observation-faithful mesh generation pipeline, and FORGE3DBench, a benchmark for evaluating observation faithfulness under realistic multi-view conditions. FORGE3D directly encodes geometric and visual observations into the generative latent space through geometry and appearance latent forcing, initializing generation from observed evidence. Applied to a pretrained generator, latent forcing already improves geometric fidelity without additional training, and finetuning with correspondence-aware multi-view attention further improves appearance. FORGE3DBench features more than 300 unique assets across 20 simulated backgrounds with 24 camera viewpoints, encompassing varied object placements, lighting conditions, and occlusions. Across three benchmarks, FORGE3D achieves state-of-the-art multi-view reconstruction performance in both geometry and appearance and improves downstream 6D pose estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.