TALE: Training-Aware Low-Rank Emulation for On-Device Adaptation
Abstract
Fine-tuning foundation models on edge devices remains challenging because training often requires 3–4 more memory than inference. Emulator-based approaches address this challenge by using a lightweight surrogate for the full model, enabling adaptation within an inference-level memory budget. Existing emulators are typically constructed to preserve the full model’s forward outputs. However, forward fidelity alone does not ensure that the resulting emulator reproduces the training signals needed for fine-tuning. To address this, we propose TALE, a Training-Aware Low-rank Emulation for memory-efficient on-device adaptation. Specifically, our theoretical analysis reveals that adapter divergence is fundamentally governed by the gradient discrepancy between the emulator and the full model. Guided by this insight, TALE optimizes a direct gradient-matching objective, extracting the compact weights in closed form via SVD. This principled approach intrinsically aligns with the full model's loss geometry, yielding an emulator that faithfully replicates the training signals essential for adaptation. Across multimodal VQA, LLM personalization, and image generation, TALE consistently outperforms forward-only emulators and achieves performance competitive with full-model fine-tuning. Furthermore, it demonstrates robust generalization capabilities to unseen users and tasks, supporting decoupled server-side generation for on-device use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.