Beyond the Estimate: Wiener-Informed Decoding for Easier Training of Pixel-Space Flow Models
Abstract
Learning velocity fields directly in pixel space is difficult and slow. We show that coarse second-order image statistics can ease this task: frozen Wiener estimates improve velocity learning, with gains extending beyond the linear predictions they explicitly provide. Moreover, we observe similar but complementary gains from frozen bases and pixel decoding. These observations motivate Wiener-informed Decoding (WiD), which uses statistical knowledge both for direct prediction and to organize decoder computation without modifying native blocks. WiD combines a trainable Wiener prediction path with coefficient interfaces that inject features into the native decoder and return its predicted coefficient corrections, with these components initialized with data statistics. All introduced parameters, including the spectrum and coordinate transforms, are trained jointly under the original pixel-space velocity objective. Across model scales and decoder architectures, WiD accelerates training and improves generation quality with modest parameter overhead. It also yields smoother gradient trajectories and remains effective with representation alignment. On ImageNet 256×256, PixelDiT-XL with WiD reaches FID 1.61 at 260 epochs, matching the baseline's reported quality at 320 epochs with less than 1% additional parameters. We also validate the early gains on text-to-image tasks. These results show that statistical denoisers can contribute beyond the estimate: the computational structure also helps pixel generation learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.