acceptodds
Under review as a conference paper at ICLR 2027

Rethinking State-Control Dynamics for Unified and Efficient Transformer Adaptation

Abstract

State-based fine-tuning reduces training memory by injecting lightweight controls into model states rather than modifying pretrained weights. However, these controls generally cannot be merged into the pretrained weights and must be evaluated at every discretization step, limiting inference throughput and practical deployment. Existing approaches incur two avoidable overheads: (1) they employ fine-grained subblock-level discretization, which requires frequent control evaluation, and (2) they instantiate separate controls across layers, forcing multiple control features to share a single routing during communication, limiting adaptation flexibility. In this work, we introduce **SURF** (**S**imilarity-Guided **U**nified **R**outing of **F**eatures), a unified state-based adaptation framework that integrates coarse-grained discretization with similarity-guided routing of shared control features. SURF composes consecutive residual subblocks into layer-wise transitions, reducing control evaluations without altering the pretrained transformations. It then groups layers with similar representations and applies sparse, token-dependent routing over a shared pool of rank-one control features, enabling feature reuse across depth while preserving input-specific adaptation. Experiments on 34 datasets and 10 pretrained models show that SURF substantially narrows the inference-efficiency gap with mergeable weight-based methods while matching or outperforming existing state-based approaches at comparable training cost. These results highlight discretization granularity and control sharing as two key design principles for efficient state-based adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.