acceptodds
Under review as a conference paper at ICLR 2027

Tiled Residual Flow: Memory-Bounded Exact-Transform Generation

Abstract

High-resolution image-space generative models are limited by activation memory that grows with image area, while latent-space models reduce memory only by relying on a learned decoder. We propose Tiled Residual Flow, a VAE-free architecture that decomposes an image into a deterministic low-resolution base and a high-resolution residual with exact additive reconstruction. A global flow models the base; a shared tile-local flow models residual tiles conditioned on crops of the upsampled base. Peak residual-stage memory is set by tile geometry rather than image size: the measured per-step training floor is 10.7x below a full-frame counterpart at 512x512, growing to 167x at 2048x2048, where full-frame training exceeds an 80 GB accelerator at batch size 2. Ablations at base resolution 128 show that removing local base conditioning nearly doubles FID, while the stitching rule, halo width, tile size and model width have no detectable effect on FID; halving the base resolution to 64 roughly halves FID at the same parameter count. On FFHQ at 512x512, with 7.4M- and 7.6M-parameter models trained on the same data, this memory bound costs no FID we can detect: the tiled model at base resolution 64 matches the best checkpoints of a full-frame flow-matching baseline whose receptive field is wider than the frame, after the same 350k full-resolution training steps (FID 56.8 against 57.8) and remains level at equal training compute, under a shared evaluation protocol and with no state-of-the-art claim.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.