GAP: Gradient-Aligned Dual Perturbation for Offsite-tuning Large Language Models
Abstract
Adapting language models across organizations sometimes requires raw data and complete original models to remain within their respective owners. Offsite-tuning addresses this situation. The model owner supplies an offsite model with a weakened emulator and trainable adapters; the data owner freezes the emulator, tunes only the adapters locally, and returns them for insertion into the complete model. The challenge is to preserve plug-in utility while maintaining an obvious performance gap over the offsite model and the tuned offsite model, which encourages the upload of the adapters to get better performance. Alignment between the emulator and the complete model is central to this trade-off. Prior approaches align representations or preserve gradients through prescribed compression operations. We introduce Gradient-Aligned dual Perturbation (GAP), a dual-perturbation approach that expands optimization beyond limited compression choices. Training-based gradient-aligned perturbation learns adversarial weight changes while matching adapter gradients at virtual gradient descent steps; compression-based gradient-aligned perturbation uses Block Influence, feature-direction change between the input and output of a layer, to select low-disruption layer for deletions. Both stages use only support set at the model owner side. Across eight benchmarks, GAP improves macro Plug-in performance over OT, the strongest baseline by 3.59% and 4.53% on Llama-3.2-3B and Llama-3.1-8B respectively with larger performance gap. We theoretically justify GAP's effectiveness and empirically demonstrate its generalization across dense and MoE backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.