acceptodds
Under review as a conference paper at ICLR 2027

AUDITING TEST-TIME TRAINING IN STREAMING 3D RECONSTRUCTION: SCALE-INVARIANT UPDATES, LARGE WEIGHT CHANGES, SMALL POSE EFFECTS

Abstract

Streaming 3D reconstruction transformers increasingly incorporate test-time training (TTT) branches whose fast weights are updated online and are intended to retain long-horizon scene information. We audit this mechanism in the released Scal3R model under its published evaluation configuration on TUM RGB-D and 7-Scenes, explicitly tracing the implemented update rule and loss and evaluating interventions as paired differences relative to variation induced by window placement. We find that the deployed TTT update, inherited from large-chunk TTT and based on Muon-style orthogonalisation, is effectively scale-invariant: changing the base learning rate by 8× alters absolute trajectory error (ATE) by at most 0.023%, while the approximately 4× variation in per-token learning rates primarily reweights update directions. Although block gradients are highly redundant, orthogonalisation substantially magnifies their differences. Replacing the full gradient with that of a single block yields gradient cosine similarities of 0.77–0.99 but orthogonalised-update similarities of only 0.54–0.81, producing final fast weights that differ by 5–16%. Despite these large parameter-space differences, ATE remains within 0.14% of the full update across all tested 300-frame windows. More broadly, the fast weights undergo substantial adaptation, corresponding to an estimated 20°–34° parameter rotation, while their effect in trajectory space is small: removing the TTT branch or freezing its fast weights changes ATE by approximately 1% or less, and blockwise adaptation without retraining degrades ATE on two of three sequences, by as much as 26%. In contrast, a larger share of test-time variability originates elsewhere in the inference pipeline. Averaging predictions across two window offsets improves ATE on every sequence of both benchmarks, while ground-truth-free selection of the alignment back-end through self-consistency provides additional, though less universal, gains, reducing mean ATE by 9.9% on TUM and 8.3% on held-out 7-Scenes. These results reveal a pronounced mismatch between adaptation in parameter space and adaptation in reconstruction accuracy, and show that inference-window and alignment choices can dominate the observable contribution of the deployed TTT memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.