Bitwise-Identical Latent Reuse for Causal Video Encoders
Abstract
Causal video encoders recompute shared work across overlapping clips, but pixel overlap alone does not justify feature reuse: independently initialized clips differ in temporal support and sampling phase. We derive compatible reuse regions at multiple encoder depths. Shallow features become reusable earlier; deeper features become reusable later but avoid more computation. ClipStore combines both within a request, preserves request-local suffix state, and retains only planned reads using bounded rolling storage. On a frozen Wan2.1 encoder, all timed output comparisons across four videos match independent encoding exactly. On the reference plan, complete construction takes less time and retains fewer feature bytes than the faster fixed-cut baseline, with a speedup over independent encoding. Request-count, length, overlap, and grouping sweeps identify where source construction is amortized. The reported gain in a 50-step text-to-video pipeline is ; encoder reuse adds no output error when combined with DiT-side caching. Decoder continuation illustrates the complementary use of retained causal state to preserve an existing stream’s outputs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.