acceptodds
Under review as a conference paper at ICLR 2027

AttentionInfluence: Training-Free Weak-to-Strong Pretraining Data Selection via Attention Head Masking

Abstract

Recently, there has been growing interest in collecting reasoning-intensive pretraining data to improve the reasoning ability of LLMs. Prior approaches rely on supervised classifiers, requiring human or LLM labeling and often introducing domain-specific biases. Since attention heads are crucial to in-context reasoning, we propose AttentionInfluence, a training-free method without supervision signal. Our approach turns a small pretrained language model into a strong data selector: we identify its retrieval heads and score each document by the loss difference incurred by masking them. We apply AttentionInfluence to a 1.3B-parameter dense model to select from the 241B-token SmolLM corpus, and mix the corpus with the selected 73B-token subset to pretrain a 7B-parameter dense model on 1T tokens under a Warmup-Stable-Decay (WSD) schedule; the baseline is identical except that its 73B subset is drawn uniformly at random, so the only variable is which tokens are selected. Improvements range from 0.8pp to 3.5pp across knowledge-intensive and reasoning-heavy benchmarks (MMLU, MMLU-Pro, SuperGPQA, GSM8K, and HumanEval), and the overall average at 1T tokens exceeds that of the FineWeb-Edu classifier, a supervised selector distilled from a 70B annotator—with no labels and no classifier training. This is an effective Weak-to-Strong scaling property: a small model improves a larger one. Code is available in an anonymous repository at https://github.com/gofornlpsota/AttentionInfluence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.