Open-Sci: Beyond Scaling with High-Density Datasets for Precise Scientific Reasoning
Abstract
Large Language Models (LLMs) have made remarkable progress on reasoning tasks such as coding and mathematics, yet their scientific reasoning ability remains significantly limited, hampered by the scarcity of high-quality scientific datasets. While LLM training has long been guided by scaling laws, attention is shifting toward data quality: high-density supervision complements scale and at the margin can outperform noisier corpora. Constructing such datasets, however, remains challenging. Existing efforts either rely on LLM-generated synthetic data prone to hallucination and logical inconsistency, or on manual curation that struggles with scalability and standardization. To overcome these hurdles, we propose PrecSci, a systematic data curation pipeline centred on precision and internal consistency. It extracts knowledge from reliable scientific documents, refines questions for completeness and precision, applies multi-stage filtering to eliminate redundancy and noise, and refines answers through a generator and a critic that scores candidates along multi-dimensional quality criteria, retaining only high-quality samples. Leveraging PrecSci, we build Open-Sci, a compact corpus of 186k scientifically precise question–answer pairs. Despite being less than one-sixth the size of the state-of-the-art MegaScience dataset ( 14.88%), Open-Sci consistently improves LLM performance, lifting Qwen3-8B by 7.08 points on average across diverse scientific benchmarks (e.g., PHYSICS, ChemBench, MMLU-Pro).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.