acceptodds
Under review as a conference paper at ICLR 2027

SwiftVR: Real-Time One-Step Generative Video Restoration

Abstract

Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusion-based VR models remain difficult to deploy on consumer-grade GPUs due to two main bottlenecks: quadratic spatial attention at high resolutions and the latency-memory overhead of large video autoencoders. We present SwiftVR, a streaming one-step generative VR framework that reduces both bottlenecks under a causal chunk-wise protocol. For attention, mask-free shifted-window self-attention uses deterministic indexing to gather spatial windows into dense tensors. This design keeps all attention operations on the dense scaled dot-product attention path, eliminating masks, cyclic shifts, padding, and hardware-specific sparse kernels. Consequently, the trained model transfers directly to consumer-grade GPUs without retraining or custom kernels. For autoencoding, a lightweight Restoration-aware Autoencoder enables fast chunk-wise decoding while preserving reconstruction quality. On a single H100, SwiftVR sustains 31 FPS at and 14 FPS at (4K), whereas all compared diffusion-based VR baselines exceed the memory limit at 4K. On a consumer-grade RTX 5090 GPU, SwiftVR reaches 26 FPS at (1080p). To our knowledge, SwiftVR is the first generative VR model to achieve real-time 1080p streaming on a consumer-grade GPU, while attaining strong no-reference perceptual quality with lower inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.