acceptodds
Under review as a conference paper at ICLR 2027

Learning Adaptive Metadata for Generative Space-Time Video Super-Resolution

Abstract

Space-time video super-resolution (STVSR) reconstructs high-resolution, high-frame-rate video from sparse, degraded pixels, but a generative receiver cannot reliably recover source-specific structure that was never transmitted. We introduce MetaST, a sender–receiver framework that learns what source information to preserve under a shared transmission budget. Pixel anchors and structural metadata compete for the same bytes, valued by their reconstruction benefit to a fixed receiver system. A lightweight Metadata Value Predictor learns these values from offline receiver responses and guides allocation using source and encoder statistics, without sender-side reconstruction. On HC-Meta, a human-centric benchmark with aligned structural representations, reallocating rate to source structure improves Average PSNR by 3.0 dB over the strongest pixel-only baseline at the lossless-anchor reference budget. On the unpruned candidate grid, learned valuation recovers 79.4% of the oracle allocation gain and reduces regret by 33% relative to a rate- only predictor. H.264/H.266 and general-video experiments further show how the useful representation depends on the relative efficiency of pixel and metadata coding. MetaST provides a modular rate–value formulation for learning what to preserve for a generative receiver.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.