acceptodds
Under review as a conference paper at ICLR 2027

KletterMix: Climbing Toward High-Quality German Pretraining Data

Abstract

High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English coun- terparts: they are often smaller, less carefully curated, weakly documented, and rarely validated through controlled training experiments. We introduce KletterMix, a high-quality German corpus for language model pretraining and annealing, de- signed as reusable dataset artifact for the natural language processing and modeling community. KletterMix is built by translating a state-of-the-art English pretraining corpus into German while preserving document boundaries, metadata, source struc- ture, and topical diversity. This construction yields a German corpus with the scale and diversity of a modern pretraining dataset, while enabling direct comparison to its English source. We document the dataset through a broad set of corpus-level analyses, including translation quality, document length distributions, topic cov- erage, and source composition. A COMETKiwi-supervised proxy enables scalable filtering, bringing the corpus’s aggregate linguistic profile closer to a matched German-web reference without uniformly selecting syntactically simpler text. Beyond dataset construction, we evaluate KletterMix through controlled pretraining and annealing ablations against established German corpora, we show that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations. These results demonstrate that carefully curated translated data can substantially strengthen the German pretraining data ecosystem.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.