SOLA-SR: Learning Scale-Ordered Latent Subspaces for Arbitrary-Scale Super-Resolution
Abstract
Latent representations have driven significant recent advances in large language models, and advance efforts to adapt such representations to computer vision have inspired new paradigms for image generation. However, discrete latent representations constrain the modeling of naturally continuous image features, especially in super-resolution (SR), where reconstruction requires richer, finer-grained image features. Inspired by COLA-DLM’s separation of continuous latent representations from conditional decoding, we propose Scale-Ordered Latent Subspaces for Arbitrary-Scale Super-Resolution (SOLA-SR). SOLA-SR separates shared image encoding from scale-dependent selective decoding, extracting image features in a single encoding pass and progressively exposing them during decoding as the target scale increases. Specifically, a learnable orthogonal transformation reorganizes feature channels while preserving their spatial layout. Continuous scale gating selects a compact subset of feature dimensions at lower scales and progressively activates additional channel groups as the scale increases. We algebraically fold the selected subspace’s linear mapping into the decoder’s first query layer, thereby skipping computations associated with inactive input dimensions. A single encoding thus supports deterministic reconstruction at multiple scales without iterative latent-variable sampling. By coupling scale-dependent representation organization with selective decoding, SOLA-SR offers a path toward combining fine-grained scale control, reconstruction quality, and inference efficiency in arbitrary-scale super-resolution. Experiments indicate an overall trend toward faster end-to-end inference and distinguish the roles of representation organization and iterative latent-variable generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.