W3D: Where, Who, and Why? Diagnosing Multi-View 3D Fusion
Abstract
Multi-view fusion assessment should explain local geometric errors rather than merely assign reliability scores. We introduce W3D, a query-centric framework that formulates fusion assessment as three coupled diagnostic objectives: Where localizes unreliable geometry, Who traces view-level influence through leave-one-view-out effects, and Why predicts multi-label observation factors. W3D reconnects each fused-surface query to its original multi-view observations using a point-normal-residual (PNR) representation that captures relative depth, viewing geometry, and surface orientation. A shared Set Transformer supports gated Where Who Why message passing, while stop-gradient operations decouple task-head updates through these messages without preventing joint learning of shared features. Supervision combines reference geometry with leave-one-view-out re-fusion, whereas inference requires only the fused surface, original observations, and camera calibration. On Skoltech3D, W3D consistently improves all six diagnostic metrics over the matched parallel Base with identical PNR inputs, remains effective across varying numbers of input views, and supports multiple fusion algorithms under the same diagnostic formulation through separate training. These results show that local fusion quality can be expressed as a traceable diagnosis linking geometric errors to their supporting observations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.