acceptodds
Under review as a conference paper at ICLR 2027

Adapter-free Parameter-efficient Fine-tuning with Tiny Routers

Abstract

Mixture-of-experts fine-tuning (MoE-PEFT) routes tokens among experts that are newly trained or that span every singular component of the pretrained weights, so it adds computation on top of the frozen backbone. Can it instead use the pretrained model’s own spectral components as experts, selecting only some of them for each token and learning how strongly to weight them? The frozen model already holds much of what adaptation needs, but a strict per-token bud- get over its spectral components does not train when candidates are ranked by the router’s gate logit, which encodes only how far a gate departs from one, not how much its group contributes. Ranking them by spectral energy makes the budget trainable (untrained, it keeps the base accuracy where random selection scores zero), and gates that can amplify the selected groups improve its accuracy. Execution Tuning builds on these findings: it decomposes each MLP weight ma- trix into frozen spectral groups that exactly reconstruct it and trains only a tiny router, 0.055% of Qwen3-8B’s parameters, that learns token-conditioned magni- tudes over an energy-selected support. Against MoE-PEFT baselines on six LLMs and six tasks, it trains 23–100× fewer parameters than MixLoRA and MoORE (with an 11–25% larger frozen bank), counts 1.5–1.7× fewer prefill FLOPs than the cheapest baseline, fewer even than the dense model, and runs faster than all of them, partly through implementation efficiency. It stays within 2 points of the strongest baseline on most model–task pairs, although MixLoRA remains more accurate on three quarters of them, and in single runs it also applies to vision– language, vision, and speech models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.