acceptodds
Under review as a conference paper at ICLR 2027

FoRA: Forward-Only Residual Adaptation for Federated Fine-Tuning of Large Language Models

Abstract

Federated fine-tuning of large language models (LLMs) faces a tension between avoiding client-side backpropagation through the LLM backbone and retaining first-order optimization for effective adaptation. Existing methods either keep trainable modules inside the backbone, preserving first-order gradients at the cost of backbone backpropagation, or adopt zeroth-order optimization for forward-only execution at the cost of gradient-estimation error. We propose FoRA, a federated fine-tuning framework that keeps the LLM backbone forward-only while preserving first-order optimization for the trainable adaptation. FoRA detaches compact multi-level representations from a frozen backbone, learns a lightweight factorized residual with a direct Split branch and a globally weighted Bridge branch, and regularizes local training against a fixed global reference in function space. Our analysis verifies the execution-graph memory separation, characterizes the local Jacobian geometry of the auxiliary parameterization, characterizes the local dynamics induced by global-reference regularization under a fixed-tangent model, and establishes non-convex convergence for the federated protocol. Empirically, FoRA reduces peak allocated client GPU memory by 27.3%-59.2% over first-order baselines while achieving comparable quality, and outperforms FedMeZO by 6.99-18.37 points in Final-10 Accuracy on classification and 4.74 points in Final-10 ROUGE-L on generation, with 65.8%-74.7% less training time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.