DDD3R: Directional Decomposition and Dampening for Recurrent 3D Reconstruction
Abstract
Recurrent 3D reconstruction models process video streams by iteratively updating a latent state, yet we find that this update process suffers from three compounding failures. Update magnitudes are systematically too large, causing progressive degradation over long sequences. All existing adaptive gating mechanisms collapse to near-constant values, providing no meaningful temporal adaptivity. And update vectors contain substantial directional redundancy, repeating historical drift rather than encoding new geometric information. These failures share one cause. Any scalar gate modulates the update uniformly across all directions in feature space, so it cannot distinguish redundant drift from novel content and reduces to blunt magnitude control. Motivated by this diagnosis, we propose DDD3R (Directional Decomposition and Dampening for Recurrent 3D Reconstruction), a training-free module that replaces scalar gating with directional decomposition. DDD3R projects each per-token state update onto a tracked drift direction and applies differentiated suppression, strongly attenuating drift while preserving orthogonal, novel information. This formulation strictly generalizes prior approaches. Constant dampening, temporal braking, and full directional decomposition all emerge as special cases of a single update rule parameterized by how much to differentiate drift from novelty. A drift-energy-adaptive extension parameterizes a continuous spectrum from full directional decomposition toward isotropic dampening that can be tuned per scene. Applied as a plug-in to existing models without retraining, a single fixed configuration (DDD3Rbrake, the scalar relaxation of the full rule) reduces camera pose error by 62-68% on long indoor sequences (TUM, ScanNet at 1000 frames), and the directional-decomposition variant (DDD3Rortho) yields a further 13% relative reduction over DDD3Rbrake on TUM where drift carries mostly harmful repetition. DDD3R also improves video depth estimation by 13-20% across three benchmarks (KITTI, Bonn, Sintel), while adding no learnable parameters and negligible computational overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.