acceptodds
Under review as a conference paper at ICLR 2027

ReasonFit: Coupling VLM Structural Reasoning with Differentiable Multi-view Fitting for Primitive-Based Reconstruction

Abstract

Recovering a part-structured 3D model from multi-view images requires deciding which parts are present and estimating their geometry. These decisions depend on each other: an omitted part can distort its neighbours, while a restrictive primitive can fragment a single component. We present ReasonFit, a closed-loop framework that couples vision-language structural reasoning with differentiable geometric fitting. Within this loop, a vision-language model (VLM) proposes part structure and qualitative shape constraints, while differentiable fitting estimates geometry from multi-view appearance and silhouettes. We use deformable SuperFrustum primitives to capture curved profiles and varying cross-sections within a single part, with VLM shape cues guiding the allowed deformations. In turn, the fitted assembly and unexplained image regions guide the VLM in distinguishing missing parts from fitting errors and revising the structure. The updated assembly is jointly refined while preserving the fitted state of retained parts. Experiments across diverse objects show that ReasonFit produces more compact reconstructions than scaffold-based fitting. Evaluations of geometric fidelity, coverage, and dedicated part recovery further reveal the trade-off between compactness and preserving distinct components.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.