acceptodds
Under review as a conference paper at ICLR 2027

DecodeAlign: Post-Decoding Representation Alignment for Video Coding for Machines

Abstract

With the increasing use of videos for machine understanding, the objectives of human-oriented video coding differ from those of machine tasks, making the improvement of machine readability for compressed videos an important challenge. However, existing approaches for machine tasks often require modifications to the coding process, new bitstream designs, or joint optimization with machine models, limiting their applicability in existing video coding ecosystems. In this paper, we revisit VCM from a representation perspective and reveal that the machine readability of compressed videos depends not only on the information preserved during encoding, but also on the effectiveness of decoded content representations for downstream machine models, which is not constrained by the original coding process. Based on this observation, we propose DecodeAlign, a post-decoding representation alignment framework for compressed video machine tasks. DecodeAlign keeps the original codecs, bitstream structures, and downstream task models unchanged, and introduces an independent representation alignment network after decoding to reduce the representation discrepancy between decoded and original content representations in the representation space of machine models, thereby improving machine task performance. Experiments demonstrate that DecodeAlign effectively improves machine task performance across eight host codecs, achieving up to 45.6% BD-rate reduction, and generalizes across codecs according to compression distortion distributions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.