When Does Adaptive Test-Time Augmentation Pay Off? A Matched Execution Study in Dense Vision
Abstract
Adaptive test-time augmentation (TTA) can allocate computation usefully and run faster without improving quality–runtime utility. We develop a comparator-aware procedure that measures complete systems and translates utility deficits into improvement requirements. A frozen SegFormer-B2 router on 400 held-out ADE20K images exhibits two effects across RTX 3090, T4 and A100 environments. At 0.8 cross-entropy (CE)/second, A100 batching increases percentage savings over multiscale from 24.7% to 28.3%, but reduces absolute savings from 19.24 to 10.12 ms/image, below the 17.48 ms required by its quality penalty. On T4, routing beats multiscale in utility but loses to native inference, the best static at that preference. The recorded batching reversal on RTX 3090 and A100 holds throughout CE/second, not only at the historical price. Removing observation/head computation alone leaves substantial deficits; a 2,000-image cross-fitted evaluation retains negative batch-8 mean surplus. Prospectively frozen ACDC calibration changes 30.0% of decisions, but runtime absorbs 85.2% of the CE gain, leaving superiority unresolved. Matched execution identifies the relevant comparators and the improvement needed to outperform them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.