acceptodds
Under review as a conference paper at ICLR 2027

Grounding Generative Verification in 3D Content

Abstract

Recent advances in multimodal large language models (MLLMs) and 3D generation have significantly expanded AI capabilities in understanding and synthesizing 3D content. However, beyond visual fidelity, task-specific correctness grounded in 3D evidence remains a major challenge. We argue that building reliable 3D systems requires not only stronger generation and reasoning, but also robust 3D verification, which offers a crucial mechanism for scaling inference by providing direct feedback for candidate selection and iterative refinement. We investigate the capabilities and supervision required for effective verification, which demands integrating multi-view partial observations and disentangling true 3D geometry from projection-dependent appearances. In this work, we make three contributions. First, we introduce 3DVerifierBench, a benchmark spanning 10 tasks that evaluates decisions, rejection explanations, and responses to controlled changes in specifications and candidates. A human study reveals a 21.6 percentage-point gap between humans and the strongest tested model, GPT-6 Astra, while explanation-aware analysis exposes further weaknesses in rejection reliability. Second, we analyze verification decisions across different fine-grained factual categories, with reviewed cases exposing incorrect factual values and target binding. Motivated by these findings, we develop a scalable data construction pipeline that links decisions to specific facts and contradictory evidence. It combines multi-view fact extraction, controlled 3D editing, and native annotations to produce a larger dataset with 14,126 samples, with explicit contradiction explanations and linked counterfactuals. Third, we examine the potential of verification under different requirements to support 3D generation through candidate selection and recursive refinement. In human-labeled sixteen-candidate pools, oracle selection improves complete factual success from 26.7% to 90.0%. In a separate eight-input study, human-guided recursive refinement satisfies all seven requirements for every input, compared with three under parallel Oracle@8. Together, our benchmark, analysis, and data pipeline establish a foundation for developing 3D verification as a capability in its own right and translating it into more reliable generative 3D systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.