ResFuser: Decomposing Residual Fusion in Aggressively Token-Dropped Vision Transformers
Abstract
Vision Transformers operating on dense scientific and medical inputs face a sequence-length bottleneck, motivating aggressive token dropping while preserving the information needed for restoration. We introduce ResFuser, a unified investigation framework that compares two existing residual topologies and four token-sampling rules with a shared backbone and controlled token and training budgets. Across five empirical cells spanning climate, medical, and natural-image restoration in deterministic and diffusion regimes, both full-resolution residual routes substantially improve over the no-residual ablation, while their relative ranking changes across cells. Within the evaluated 7.4M-parameter backbone, the better topology in each cell gives validation RMSE within % of a dense ViT/DiT baseline at 42-59% fewer analytical forward FLOPs. The tested sampling rules produce smaller changes than residual design; masked-region analysis also shows that input salience need not identify the regions where restoration quality improves. These findings motivate a staged design procedure: establish a residual route, compare topologies under the target budget, and then evaluate task-aligned sampling. ResFuser decomposes these choices for dense restoration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.