acceptodds
Under review as a conference paper at ICLR 2027

Beyond What to Select: Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training

Abstract

Data selection accelerates training by prioritizing informative training samples, yet existing methods primarily optimize *which samples to select* while keeping *how much data to use* fixed throughout training. As a result, the temporal role of data volume remains largely underexplored, even in dynamic selection. In this work, we find that the instantaneous data volume directly shapes subset-based optimization dynamics, revealing a regularization–fidelity trade-off: lower ratios induce stronger stochastic regularization while reducing computation, whereas higher ratios provide more faithful and stable updates. Therefore, fixing data volume throughout training constrains optimization and can lead to suboptimal performance under a given training budget, motivating a rethink of data selection beyond fixed-ratio training. Building on this insight, we propose **PODS**, a **P**lug-and-play **O**scillatory **D**ata-volume **S**cheduling framework that alternates low- and high-ratio phases while preserving a prescribed cumulative data budget. Given a target ratio, the schedule is analytically determined without dataset-specific tuning and can be directly integrated with existing static or dynamic selectors. Experimental results across diverse vision and language tasks show that **PODS** consistently extends the efficiency frontier of fixed-ratio selection, reducing ImageNet-1k training cost by 40% while improving accuracy by 0.4%, and achieving approximately 2× faster LLM instruction tuning without performance degradation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.