acceptodds
Under review as a conference paper at ICLR 2027

ZERO-TRAINING SURFACE STATISTICS FOR PRETRAINING-CORPUS FILTERING

Abstract

The value of pretraining corpora depends not only on their size, but also on how effectively low-quality content is filtered. Heuristic filters are efficient, interpretable, and scalable, but their hand-designed rules are pipeline-specific and do not generalize across languages. We extract the underlying measurements from seven widely used pipelines and summarize the corpus problems they target in an operational taxonomy of seven corpus-quality failure modes. Building on heuristic filters' view of language structure, we form ZenoFilter from how these measurements respond across failure modes, yielding a training-free adaptive detector that measures departures from ordinary text in redundancy and lexical composition and supports an adjustable corpus-level operating point. On real WARC documents, ZenoFilter outperforms the filtering layers of four widely used pipelines: Gopher/C4, FineWeb, Dolma, and RefinedWeb. At matched keep rates on symmetric English, it leaves 2.16–2.63× fewer document-level and 1.74–2.12× fewer token-level residual artifacts while running about six times faster. Against the closest adaptive comparator DCAD, ZenoFilter is never worse across its keep-rate curve. On non-English monolingual text it performs comparably to FineWeb 2 and DCAD, the only comparators that cover it. It processes multilingual and code-switched documents without language gates or language-specific thresholds, and discriminates mixed-language documents comparably to English ones. In downstream validation it achieves lower held-out loss than FineWeb across the six content categories and on the source-aligned Nemotron-CC validation set; the FineWeb advantage persists on the equal-weight mean of the externally defined Pile-22 suite.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.