acceptodds
Under review as a conference paper at ICLR 2027

GRIP: Geometric Refinement and Adaptive Information Potential for Data Efficiency

Abstract

Efficient language-model pre-training requires deciding how to divide a fixed token budget across semantic regions and which examples to retain within each region. We introduce **GRIP** (Geometric Refinement and Adaptive Information Potential), which combines model-conditioned allocation with length-aware geometric selection. At the cluster level, static quality and normalized held-out loss reduction determine token budgets. Within clusters, a length-rectified inverse-density sampler increases geometric coverage while retaining long documents. We evaluate **8B and 16B Mixture-of-Experts models** on **eight code benchmarks** with budgets up to **300B tokens**. Across **five paired seeds at 300B tokens**, GRIP exceeds UniGeM by **0.9 and 1.3 average-score points** at 8B and 16B, respectively. Across **three paired seeds at 100B tokens**, GRIP reaches **35.5 ± 0.3**, **1.2 points above UniGeM**. The ablation separates the gains from cluster allocation and document selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.