SparseTransfer: Does Calibration-Free Sparse Attention Generalize Across Domains?
Abstract
Training-free sparse attention methods are increasingly evaluated as if they were universal: a mechanism validated in one domain is often assumed to transfer to others without recalibration. We test this assumption directly by porting two independently-published, training-free block selection mechanisms EntropyInfer, an entropy-guided adaptive budgeting method, and a cumulative-energy-based thresholding method into a shared attention framework (XAttention), configuring each to match XAttention's own tuned baseline operating density, and evaluating both against that baseline across two domains: long-context language modeling (RULER, Llama-3.1-8B, context lengths 4K–128K) and long-video question answering (Video-MME, Qwen2-VL-7B, 900 videos). We observe a striking, consistent asymmetry: both mechanisms transfer far more faithfully to video-understanding (97–101 of baseline accuracy) than to language (77–99.6 of baseline, with the largest gap at the longest context length), and this pattern holds across both tested mechanisms, with the cumulative-energy method consistently more robust than EntropyInfer in both domains. We verify our findings are not artifacts of implementation error through multiple independent checks, including direct inspection of the patched selection function and measured selection density. Our results suggest that "calibration-free" claims in the sparse attention literature are domain-contingent rather than universal, and that block-selection quality, not merely computational budget is a meaningful, underexplored axis of transfer failure
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.