acceptodds
Under review as a conference paper at ICLR 2027

When Seeds Change the Evidence: Seed-Conditional PEFT Comparisons

Abstract

Average parameter-efficient fine-tuning (PEFT) performance can appear stable while individual stochastic runs exhibit severe lower-tail failures. In a matched LLaMA-3-8B panel spanning 20 seeds, eight commonsense tasks, and rank-8 LoRA, DoRA, and rsLoRA, DoRA averages 84.65% but reaches 62.17% in its worst seed-task cell; rsLoRA averages 83.55% but reaches 52.47%. Lower-tail reliability is only one part of the evaluation problem: even when the directional winner is stable, the magnitude of a reported advantage can depend on the selected seeds. We separate these two aspects using complementary finite-panel summaries. The Seed Recovery Profile (SRP) describes method-level worst-task recovery, while the Claim Survival Profile (CSP) conditions on -seed subsets supporting a favorable practical claim and measures how often the complementary seeds retain the same margin. For DoRA over LoRA, the direction persists on all 13 complements of favorable one-seed reports at , but a 0.5-point margin persists on only 3 of 5. Exhaustive enumeration shows that small reporting budgets can estimate mean accuracy closely while missing lower-tail outcomes and overstating practical margins. A prospectively frozen five-seed, four-task Qwen3-8B panel documents severe seed-indexed differences on a second backbone. Complete paired outcomes remain the primary empirical record; SRP and CSP expose two aspects that average accuracy omits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.