acceptodds
Under review as a conference paper at ICLR 2027

Learning Selective Reference Dynamics for Perceptual Video Compression

Abstract

Recurrent video codecs must exploit decoded history while limiting the reuse of reference errors. We introduce **SRDVC**, which separates predictive reference evolution from reference-conditioned coding. A **Latent Dynamical System (LDS)** factors reference prediction into dynamics estimation, state evolution, and observation prediction. **Dynamical Manifold Alignment (DMA)** uses current-step features to query the predicted reference, while residual paths convey content not captured by the prediction. A local perturbation model motivates direction-dependent propagation: under bounded total input, uniform contraction of an orthogonal component bounds its accumulation. This conditional analysis provides a design rationale, while temporal-subspace probes characterize the behavior learned by the codec. At matched compute on **UVG**, replacing LDS, DMA, or both degrades perceptual rate–distortion performance; replacing both increases **LPIPS/DISTS BD-Rate by 52.41%/65.76%**. Across these variants, stronger relative directional selectivity accompanies better compression. On **UVG, MCL-JCV, and HEVC-B**, SRDVC achieves mean BD-Rate reductions of **72.67% under LPIPS** and **63.97% under DISTS** relative to **GLC-video**, with efficient steady-state 1080p P-frame coding and favorable viewer preferences over the tested subjective-study baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.