WindowFold: Folding Sliding-Window Inference into a Single Forward Pass for Dense CLIP
Abstract
Every training-free method that turns CLIP into a dense predictor is evaluated with sliding-window inference: the image is cut into overlapping crops at the resolution CLIP was pretrained on, each crop is encoded separately, and the predictions are averaged, so the encoder processes the image area four times over. We show the window is a workaround for a train–test mismatch, and that it is doing two separable things. First, attention trained to spread its mass over 196 keys misallocates it over more: a plain high-resolution forward collapses while windows over the same pixels hold up. Second, a 224 crop silently caps how far a surgery may mix features, a cap worth 8.9 mIoU to a single pass on Cityscapes that no reweighting of attention can express. We replace each with a learned correction fit to the window's own output, with no labels, text or external data: a rank-16 LoRA on the attention projections, and a Gaussian footprint on the last block's mixing whose two to six widths are fit by the same distillation. Across three surgeries and four benchmarks, one forward matches or exceeds sliding-window inference in all 12 settings under the published protocol at – fewer FLOPs. On SC-CLIP, the training-free state of the art, one forward through the full method is above the published numbers on four of five benchmarks and on their average (42.68 against 42.4) with a single configuration. The fitted footprints are themselves a finding: five surgeries trained independently place the attention footprint at 5–7 patches, half a window, fixed in patch units rather than scaling with the image; across four backbones it follows the pretraining crop, not a pixel extent, and on L/14 one forward beats the ensemble on all five benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.