Per-Image Latent Optimization at Extreme Bitrates
Abstract
Neural codecs compress with an amortized encoder: a single network, trained to perform well on average, encodes every image in one forward pass. Per-image latent rate–distortion optimization (RDO) instead refines the transmitted latent for each individual image at encoding time. It is commonly regarded as a marginal refinement, worth a few tenths of a dB at conventional rates, but that assessment was formed when bits were plentiful. Generative decoders now sustain reconstruction at bitrates orders of magnitude lower, where every bit the encoder misallocates is a visible fraction of the budget. We show that in this regime the value of per-image optimization is not a constant of the technique but grows systematically as the operating rate falls. We measure the gain as a rate-equivalent margin, the additional bits the amortized encoder requires on the same image's own rate ladder to match the optimized quality. Across three architecturally distinct codecs whose ladders span three orders of magnitude in rate, the median margin rises as rate decreases and saturates at a backbone-specific ceiling. At the lowest rates, the amortized encoder needs to times its bits to match per-image optimization, corresponding to median per-image BD-rate savings of –. A continuous-relaxation analysis traces this growth to two coupled trends: the per-image suboptimality of the amortized encoder widens steadily as rate falls, and the share of it that discretization-aware optimization captures becomes meaningful only at the lowest rates. The gains transfer to held-out perceptual metrics the optimizer never saw, precisely in the low-rate regime; optimizing a mismatched objective instead erases or reverses them. Practically, the gain saturates in compute: doubling the standard optimization budget adds at most a relative , so a modest budget suffices whenever content is encoded once and decoded many times. Together, these results identify the amortized encoder, rather than the decoder, as a principal bottleneck of neural compression at extreme rates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.