PAIRNeRV: Patch-Owned Appearance and Innovation for Neural Video Representation
Abstract
Neural video representations offer compact reconstruction, but fitting them separately to each video incurs substantial encoding cost. Feed-forward token-based methods avoid this optimization by inferring video codes for a shared decoder. However, using a global token memory leaves temporal sharing and spatial relevance implicit in the stored representation, while reconstruction repeatedly retrieves content through attention. We propose PAIRNeRV, a feed-forward framework that organizes video state by temporal role and spatial ownership. A clip-wide encoder gathers global context, from which two writers produce appearance codes shared across frames and innovation codes specific to individual frames, both assigned to spatial patches. A shared decoder reconstructs requested patches through direct lookup of their appearance and innovation codes, without global retrieval attention. Appearance and geometry projections are reused across frames. Eight-bit storage complements the design by retaining more code values within a fixed bit budget. Experiments show improved reconstruction over the compared feed-forward baselines under a common video-state budget and transfer to unseen datasets without fine-tuning. Comparisons with global-attention readout demonstrate faster resident-state decoding, while controlled studies examine the roles of state organization, storage allocation, and cross-frame reuse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.