acceptodds
Under review as a conference paper at ICLR 2027

COMPUTATIONAL DEPTH AS AN ADAPTATION VARIABLE FOR PARAMETER EFFICIENT FINE TUNING

Abstract

Parameter efficient fine tuning adapts pretrained language models with a small number of trainable parameters. Most existing methods seek greater adaptation capacity by changing how these parameters modify a fixed pretrained computation graph. We study a complementary possibility in which the amount of pretrained computation itself becomes part of adaptation. We introduce recurrent PEFT by repeatedly executing a frozen pretrained middle stack while training lightweight adapters around the reused computation. Building on residual damped recurrence introduced in prior work, we consider both iteration specific and iteration shared adaptation. This construction increases effective computational depth without adding backbone parameters and enables direct comparison with conventional single pass PEFT under matched trainable parameter budgets. Across Qwen and Llama backbones from 0.6B to 8B parameters, recurrent adaptation remains broadly competitive on general reasoning and shows gains on GSM8K across all evaluated backbones. Minerva reveals a more pronounced backbone-dependent effect, with substantially larger gains on Qwen than on Llama models. We further find that these effects depend on the underlying PEFT parameterization and that the choice of recurrent execution remains consequential after fine-tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.