Coded Computation: Perceptual-Semantic Acceleration for Video Diffusion
Abstract
Efficient video diffusion is commonly approached by making generative computation cheaper through fewer denoising steps, quantized operators, feature reuse, or selective computation. We explore a different perspective—can computation itself be coded and subsequently reconstructed? Inspired by coded acquisition in computational imaging, we introduce coded computation, a new formulation that jointly designs which generative computations are explicitly performed and how the omitted information is recovered. We instantiate this principle with a perceptual-semantic computational encoder and a lightweight code-aware generative decoder. The encoder combines bottom-up visual cues with top-down semantic relevance from the text prompt to construct a time-varying compute code, allocating richer computation to perceptually and semantically important content while encoding less critical content more efficiently. The decoder then leverages the compute code and strong spatiotemporal and diffusion priors to reconstruct the omitted generative information. Extensive experiments demonstrate that coded computation substantially accelerates video diffusion inference while preserving visual quality, temporal consistency, and prompt fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.