acceptodds
Under review as a conference paper at ICLR 2027

Test-Time Scaling for Multi-View Diffusion Models: Generation, Verification, and Search

Abstract

Test-time scaling (TTS) improves generative performance by allocating additional inference computation, yet its effectiveness for structured outputs such as multi-view diffusion models remains underexplored. In this work, we systematically investigate how generation, verification, and search interact when the generated views must remain consistent. We first find that joint multi-view generation exhibits stronger Oracle scaling than independent per-view generation across multiple evaluated models, while revealing a fundamental tension between preserving cross-view consistency through set-level selection and improving individual views through view-level selection. To address this challenge, we introduce DINO-centroid, a lightweight verifier that leverages cross-sample consensus for complete-set selection and uncertain-view identification. Our analysis shows that DINO-centroid becomes increasingly aligned with the ground-truth representation as the sampling budget grows, suggesting that the verifier itself exhibits a form of scaling behavior absent from conventional TTS verifiers. Building on these findings, we propose Exploration–Refinement Search (E-R Search), a strategy that combines set-level exploration with targeted view-level regeneration by formulating the latter as a diffusion inverse problem. Under matched inference budgets, our strategy achieves higher reconstruction quality than existing test-time search baselines. Our findings highlight the importance of accounting for component dependencies in generation, verification, and search for effective TTS of structured outputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.