acceptodds
Under review as a conference paper at ICLR 2027

SparseTransfer: Does Calibration-Free Sparse Attention Generalize Across Domains?

Abstract

Training-free sparse attention methods are increasingly evaluated as if they were universal: a mechanism validated in one domain is often assumed to transfer to others without recalibration. We test this assumption directly by porting two independently-published, training-free block selection mechanisms EntropyInfer, an entropy-guided adaptive budgeting method, and a cumulative-energy-based thresholding method into a shared attention framework (XAttention), configuring each to match XAttention's own tuned baseline operating density, and evaluating both against that baseline across two domains: long-context language modeling (RULER, Llama-3.1-8B, context lengths 4K–128K) and long-video question answering (Video-MME, Qwen2-VL-7B, 900 videos). We observe a striking, consistent asymmetry: both mechanisms transfer far more faithfully to video-understanding (97–101 of baseline accuracy) than to language (77–99.6 of baseline, with the largest gap at the longest context length), and this pattern holds across both tested mechanisms, with the cumulative-energy method consistently more robust than EntropyInfer in both domains. We verify our findings are not artifacts of implementation error through multiple independent checks, including direct inspection of the patched selection function and measured selection density. Our results suggest that "calibration-free" claims in the sparse attention literature are domain-contingent rather than universal, and that block-selection quality, not merely computational budget is a meaningful, underexplored axis of transfer failure

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.