acceptodds
Under review as a conference paper at ICLR 2027

Video Generation Models are Zero-Shot Low-Level Vision Solvers

Abstract

Video generation models learn rich priors over appearance and motion, enabling instruction-guided synthesis and editing. Here we investigate a central question: can these priors also support low-level vision without task-specific fine-tuning? We show that a single frozen video generator can serve as a zero-shot solver for a spectrum of low-level video tasks, including blind video super-resolution, colorization, archival film restoration, and restoration of AI-generated video. Our framework, Gen4LV, casts these tasks as reference-conditioned generation with complementary visual cues. That is, the degraded video guides motion, layout, and content, while a single frame restored by an image editing expert guides appearance. However, this conditioning scheme enables the generator to synthesize detailed video but does not, by itself, ensure fidelity to the input. For tasks requiring high fidelity, we further introduce frequency-gated data consistency to constrain the output using reliable observation information.In blind super-resolution, for instance, the gate restricts observation-based corrections to low frequencies, improving structural alignment while preserving generated texture and enabling inference-time control over the perception–fidelity trade-off. Across eight benchmarks spanning four tasks, Gen4LV delivers strong no-reference video quality, matching or surpassing specialized models in multiple restoration settings without task-specific fine-tuning. These results demonstrate that pretrained video generators can serve as generalists for low-level vision, pointing toward a future in which diverse visual tasks are addressed by specifying instructions, references, and constraints at inference time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.