Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning
Abstract
Multi-task supervised fine-tuning (SFT) often casts a heterogeneous data mix-ture as a single optimization problem, even though different tasks may reach their best generalization at different times. MSFT is a scheduler that exposes this mismatch through task-wise roll-out, exclusion, and rollback. However, it materializes its scheduler state as full-model checkpoints, which makes stage transitions costly to store, restore, and deploy. This paper introduces AuroSFT, a parameter-efficient framework that recasts the carried state of overfitting-aware multi-taskSFT as a compact, mergeable adapter state. AuroSFT freezes the pretrained backbone, trains only injected adapters, and rolls back adapter checkpoints at task-wise peaks. At the layer level, each adapter applies an AuroRA-inspired adap-tive nonlinear layer (ANL) to a low-rank weight factor rather than to the sample representation. The resulting update remains linear in the input, rank-bounded, and exactly mergeable into the frozen projection. On the same ten benchmarks and five backbones as the MSFT reference, AuroSFT achieves 61.36% average accuracy, compared with 59.85% for the full-model reference and 61.28%for a scheduler-matched linear-LoRA baseline. Its accuracy exceeds the full-model reference on all five backbones. A stage-log diagnosis further traces themathematics-group regression to interference that the eviction rule misreads as overfitting, and an interference-aware repair recovers this group above the full-model reference on Qwen2.5-3B. Our code is available at the anonymous repository: https://anonymous.4open.science/r/AuroSFT-80D1.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.