FinerSR: Restoring Small but Semantically Critical Regions in Generative Video Super-Resolution
Abstract
Generative video super-resolution has made remarkable progress in restoring realistic visual content from severely degraded videos. However, high global perceptual quality does not imply semantic fidelity in small but semantically critical regions, where visually plausible reconstructions may contain incorrect text or altered facial identities. To address this challenge, we propose FinerSR, a unified framework that strengthens small but semantically critical regions in both representation and supervision levels while dynamically coordinating local semantic fidelity with global perceptual quality. We first establish a strong four-step general restorer through supervised fine-tuning, on-policy distillation, and perceptual and reconstruction refinement. Building on this foundation, we introduce object-aware conditioning and region-normalized supervision to enhance the representation and optimization of small but semantically critical regions. We further develop discriminative objectives for two distinct semantic properties: alignment-aware identity and appearance matching for faces, and complete-string likelihood with edit-graph probability contrasts for text. Finally, we formulate semantic refinement as a quality-constrained joint optimization problem, using adaptive primal-dual updates to coordinate local semantic objectives with full-frame fidelity and perceptual constraints relative to the general restorer. Experiments on faces and text demonstrate that FinerSR enhances semantic fidelity in small but semantically critical regions while preserving overall restoration quality, bridging the gap between global visual quality and local semantic correctness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.