acceptodds
Under review as a conference paper at ICLR 2027

StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression

Abstract

Reconstructing dense 3D geometry from continuous video is severely bottlenecked by the unbounded growth of streaming KV caches. We present StreamCacheVGGT, a constant-budget visual geometry transformer equipped with two training-free modules. Cross-Layer Consistency-Enhanced Scoring (CLCES) tracks rank stability across Transformer layers to yield robust token importance scores and uncertainty estimates. Uncertainty-Gated Deferred Compaction (UGDC) then routes candidates through either immediate retain, merge, and evict compaction or a small deferred pool for short-term verification. This deferred pool operates strictly within the fixed per-layer budget (), ensuring the persistent memory footprint remains completely unchanged. Across diverse indoor, outdoor, long-sequence, and depth benchmarks, our method significantly improves geometric completeness and long-horizon stability while transparently reporting the expected throughput trade-off of deferred verification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.