Budget-Constrained Decomposition of LLMs
Abstract
Adapting large language models (LLMs) is highly resource-demanding: for low-resource research institutions, a full LLM may not even fit into one GPU, and federated fine-tuning still requires every client to run the full model. Existing remedies fall short. First, structured pruning selects one sub-network on a generic calibration corpus and deploys it to every client. Second, federated methods that derive per-client sub-networks bound parameters or memory, not per-token compute. To overcome this, we propose Federated Budgeted Decomposition (FedBUD), which decomposes one pretrained LLM into smaller per-client structural specialists under a hard per-token FLOPs budget. A graph-based reinforcement-learning policy selects each client's blocks, attention heads and feed-forward neurons from its architecture graph and importance statistics. The specialists are then co-trained via federated learning, with LoRA factors aggregated in the coordinates of the source model; an optional shared graph neural network modulates each specialist by its architecture graph, with representation discrepancy bounded by architecture-graph distance for its attention instantiation. FLOPs-matched experiments on six federations show that co-training outperforms independent adaptation across the six datasets, significantly on three, and that sharing structure across clients, selected on validation data, significantly improves both tasks where the selected arm was run at three seeds, whereas the learned allocator does not outperform hand-designed allocation rules. Compute-limited clients can thus co-train LLM specialists without pooling data. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.