acceptodds
Under review as a conference paper at ICLR 2027

Less Text, More Evidence: Static Pruning and the Coverage Paradox in Language Models

Abstract

Human skim reading suggests that linguistic information is distributed non-uniformly. We ask whether Large Language Models (LLMs) exhibit an analogous, statically identifiable selectivity: which linguistic structures in the input are load-bearing, and which can be removed before tokenization? We apply 15 fixed interventions-11 linguistically motivated pruning rules and four random token-deletion controls-to raw text, without LLM-based scoring or task-specific training, and evaluate them unchanged across nine LLMs and six benchmarks. Retaining only the subject-verb-object core of each sentence discards 68% of input tokens on average while retaining 63% of baseline quality; after adjusting for reduction rate, this rule retains the most quality among the linguistic rules on every evaluated model, whereas frequency-based filtering falls below the trend. The effect depends strongly on the task and the context regime. When inputs exceed the context window, pruning also changes which evidence reaches the model: on the longest LongBench-v2 documents, the pruning penalty observed on short documents disappears, and on identical items, the size of this reversal tracks how often each model must truncate the unpruned input (Spearman across models). We characterize this reversal as the coverage paradox and show that it depends jointly on recovered coverage and retained linguistic content. Our results provide a controlled map of which structures are removable and how these sensitivities transfer across models, tokenizers, and languages, and they imply that input reduction must be evaluated jointly with the context budget and truncation protocol, not by token count alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.