Tail-Calibrated Loss-Distribution Shaping for Reasoning Fine-Tuning
Abstract
Curricula based on sample ordering or selection regulate which examples are presented during training, but do not directly specify the relative weighting of token losses within each response. We propose Tail-Calibrated Loss-Distribution Shaping (TCLD), which adapts response-token loss aggregation using feedback from the current training-loss distribution. TCLD replaces arithmetic averaging with a power-mean objective whose parameter controls relative token gradient coefficients and recovers standard cross-entropy (CE) at . A Hill-tail controller calibrates at observation-window boundaries using no-gradient loss observations, reducing the relative coefficients of higher-loss tokens without a downstream-validation-accuracy search over . On a heterogeneous five-source reasoning pool, TCLD improves Macro- and Micro-averaged maj@3 accuracy over random-shuffle training (RS) by 2.74 and 2.02 percentage points, respectively, across three matched seeds with Qwen3-1.7B-Base, and outperforms the evaluated sample-level curriculum baselines in aggregate accuracy. At the 25% optimizer-step checkpoint, TCLD exceeds the full-budget RS baseline in mean Macro accuracy. Source-wise diagnostics show heterogeneous sensitivity to fixed powers and the competitive performance of automatic calibration. Gains over RS also extend to Qwen3-0.6B and SmolLM2-1.7B. Single-seed experiments on six ordered two-source streams show order-dependent gains and per-source trade-offs relative to RS. TCLD requires no changes to the model architecture or inference procedure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.