acceptodds
Under review as a conference paper at ICLR 2027

Toward AI-Friendly Data Quality Metrics: A Feature Density Perspective

Abstract

In Large Language Models (LLMs) fine-tuning stage, input sequences differ in the proportion of task relevant tokens that help the model predict target sequences, and quantifying this variation can help clarify the relationship between data and fine-tuning performance. Therefore, we introduce Feature Density: the fraction of input tokens whose deletion increases a pretrained model’s loss on the target sequence. We validate the tokens it identifies by comparing them to known feature tokens in synthetic data and human-annotated feature tokens in real world data. In synthetic data, Feature Density increases with the proportions of task relevant tokens. Across twelve real-world datasets using Qwen LLMs, Feature Density varies by task, with higher densities in translation and math reasoning. In both settings, Adam, AdamW, Muon, and SGD achieve lower test loss when fine tune on higher density data within the same tasks. These results support Feature Density as a metric of task relevant token proportions and its relation with model fine-tuning performance across different optimizers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.