acceptodds
Under review as a conference paper at ICLR 2027

STELLA-Flow: Dynamic Data Recipe Valuation for Data-Efficient Model Iteration

Abstract

As large language models mature, development is shifting from a model-centric to a data-centric paradigm. Model capabilities depend not only on data scale, but also on data quality, diversity, and composition. However, existing studies largely focus on individual-sample valuation or aggregate metric optimization, leaving the dynamic value of complete data recipes under evolving model feedback insufficiently characterized. We propose STELLA-Flow, a dynamic framework that evaluates candidate data through intrinsic quality, conditional model gain, and business value. It first assesses data quality, then uses a small number of real model updates to predict which existing errors candidate data recipes would correct and which new errors they would introduce, and finally incorporates scenario importance and error costs to select data for the current model. Across model scales, STELLA-Flow outperforms random selection and static quality filtering. Under a unified advertising content moderation benchmark, a 2B model updated with STELLA-Flow-selected data outperforms a 32B model updated with randomly selected data, GLM-5.3, and DeepSeek-V4-Pro. These results demonstrate that STELLA-Flow provides a unified method for dynamic data valuation and selection across model scales, enabling limited training budgets to continuously shift toward higher-utility, lower-risk data recipes as model capabilities evolve. This improves the data efficiency of iterative model development and offers a new path for smaller models to achieve competitive business performance at lower cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.