acceptodds
Under review as a conference paper at ICLR 2027

Breaking the Memory Barrier for On-Device LLM Adaptation via Dynamic Layer Pruning

Abstract

Federated fine-tuning holds the promise of privacy-preserving Large Language Model (LLM) adaptation, yet the prohibitive memory costs limit participation from resource-constrained devices. In this paper, we propose FedPruner, a novel paradigm that enables efficient training via intelligent, on-the-fly layer pruning. FedPruner flexibly prunes the global model, creating personalized submodels based on device memory constraints. It employs a macro-micro synergistic pruning framework: a macro-level Functionality-Driven Layer Orchestration mechanism groups layers, while a micro-level Importance-Aware Layer Selection strategy prunes within groups to build device-specific submodels. We further introduce a fine-grained variant that independently prunes Multi-Head Attention and Feed-Forward Network components to precisely preserve critical architectural elements. Extensive experiments demonstrate that FedPruner significantly outperforms state-of-the-art methods with average accuracy gains of up to 11.11%. Moreover, it maintains strong robustness under varying memory constraints, yielding a 1.98% average performance improvement while reducing peak memory usage by 75%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.