acceptodds
Under review as a conference paper at ICLR 2027

SmartBCD: Smart Block Coordinate Descent via Cost-Aware Scheduling

Abstract

Full-parameter fine-tuning (FFT) of large language models requires substantial GPU memory. Block coordinate descent (BCD)-based methods address this challenge through block-wise parameter updates rather than simultaneous updates to all parameters, enabling more memory-efficient FFT. However, existing BCD-based methods either follow predetermined or randomized block schedules, or prioritize blocks primarily based on optimization signals, without explicitly accounting for the joint variation in optimization gain and computation cost across training stages. This can lead to suboptimal use of limited training resources. We propose SmartBCD, a cost-aware block scheduling method that improves resource utilization. SmartBCD estimates each block’s utility as the predicted loss reduction per unit of training time and uses this estimate to schedule updates as training progresses. A maximum-wait constraint ensures that every block continues to receive updates. We conduct experiments with TinyLlama-1.1B and Qwen2.5-7B on instruction-following and mathematical reasoning tasks. On our Qwen2.5-7B MathInstruct setting, under the same optimization-step budget, SmartBCD reduces end-to-end training time by 4.4% and final validation loss by 1.3% relative to BAdam-Random, without negligible additional GPU-memory overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.