acceptodds
Under review as a conference paper at ICLR 2027

FineWeb2-HQ-QA: Native Document-Grounded QA for Multilingual Pre-Training

Abstract

Recent advances in data curation have substantially improved the quality of multilingual web corpora, yet large-scale synthetic resources for non-English pre-training remain scarce and typically rely on machine translation from English-source data. We introduce *FineWeb2-HQ-QA*, a multilingual dataset containing 193B tokens of source documents and generated QA across 20 languages. We generate QA pairs directly in each target language from native FineWeb2-HQ web documents rather than from English-source text (*native QA*). This approach is designed to reduce dependence on English-source topic distributions by grounding generation in target-language web content. At 3B and 8B scales, we compare our native document-grounded QA against non-QA baselines and Nemotron's English and machine-translated QA. Relative to matched non-QA baselines, our dataset improves Global and Local Knowledge across all four training settings while largely preserving or improving aggregate general language understanding. Among individual QA sources, translated QA yields the largest observed Global Knowledge gains across all four settings, whereas native QA consistently shows higher observed scores than translated QA on general language understanding and achieves strong Local Knowledge gains at 8B. In natural-mixture settings, combining sources yields a more balanced profile across categories without consistently exceeding the best individual source per category. We release our dataset alongside our generation and filtering pipelines to support future work on multilingual synthetic pre-training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.