Attention Residuals for Pixel-Space Flow Transformers: An Empirical Study
Abstract
Attention Residuals (AttnRes) replace fixed residual accumulation with content- dependent retrieval over earlier residual increments. We study the published Full AttnRes operator in class-conditional, pixel-space JiT flow Transformers under matched backbone and training settings. On ImageNet 256×256, Full AttnRes reduces six-encoder FDr6 from 15.80 to 14.72 at B/16 (6.8%) and from 11.92 to 10.81 at L/16 (9.3%), with lower ratios in all six representation spaces at both scales. Learned parameters increase by approximately 0.01% and profiled infer- ence FLOPs by below 1%, although full-history storage and mixing remain im- portant systems costs. We also insert zero-query Full AttnRes into a shared 100- epoch standard-residual checkpoint. Equal-budget continuation raises IS from 224.6 for unchanged continuation to 230.6 after conversion. We explain approx- imate function preservation under explicit pre-normalization assumptions, local- ize most of an ordered ablation’s endpoint gap to condition-token history, and observe flow-time-dependent depth routing. These results support studying Full AttnRes both from scratch and as a continued-training feature, while highlighting representation-dependent effect sizes and the need for repeated-run and systems evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.