JRCC: Joint Residue and Conditional Coding for Overfitted Neural Video Compression
Abstract
Residue coding has long been one of the most widely used paradigms for exploiting inter-frame correlation in both conventional and learned video coding. More recently, conditional coding has emerged as a more efficient paradigm for exploiting inter-frame correlation in offline-learned video coding. This raises a natural question: can conditional coding also be effectively used in overfitted learned video coding? The key difference between overfitted and offline-learned video coding is that, in the former, the neural network model parameters must be transmitted. Therefore, using a large model to exploit conditional dependencies between neighboring frames is less suitable for the overfitted setting. In this paper, we find that directly applying conditional coding degrades rate–distortion performance under this overfitted setting, while a moderate increase in conditioning capacity can partially alleviate the degradation, but further scaling alone does not consistently resolve it. We therefore propose a joint residue-conditional coding framework that retains explicit motion-compensated prediction while using motion-aligned temporal features to condition residue synthesis. Experiments on multiple benchmark datasets under random-access (RA) and low-delay (LD) configurations show that the proposed framework consistently improves Cool-Chic 4.1, achieving average BD-rate reductions of 9.7% and 27.8%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.