FreeSolo: Linear-time Sparse Attention with Zero-Shot Length Extrapolation
Abstract
Linear-time sequence models typically store long-context information in a recurrent state. A fixed-size state limits how much information can be stored, while hybrid architectures that add full attention for better recall incur quadratic sequence cost. Alternatively, sparse attention with continued pretraining (CPT) can extend the effective context length. Generalization beyond the trained length remains challenging, however, and CPT can be costly. We introduce FreeSolo, a native sparse attention mechanism with a short-range sliding window branch and a long-range branch that selects keypoint cells for direct token attention. Its linear-time variant, FS-L, learns to prune the prefix before query-dependent selection through short CPT. FreeSolo demonstrates better zero-shot length extrapolation than the baselines without long-context fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.