acceptodds
Under review as a conference paper at ICLR 2027

TALE: Training-Aware Low-Rank Emulation for On-Device Adaptation

Abstract

Fine-tuning foundation models on edge devices remains challenging because training often requires 3–4 more memory than inference. Emulator-based approaches address this challenge by using a lightweight surrogate for the full model, enabling adaptation within an inference-level memory budget. Existing emulators are typically constructed to preserve the full model’s forward outputs. However, forward fidelity alone does not ensure that the resulting emulator reproduces the training signals needed for fine-tuning. To address this, we propose TALE, a Training-Aware Low-rank Emulation for memory-efficient on-device adaptation. Specifically, our theoretical analysis reveals that adapter divergence is fundamentally governed by the gradient discrepancy between the emulator and the full model. Guided by this insight, TALE optimizes a direct gradient-matching objective, extracting the compact weights in closed form via SVD. This principled approach intrinsically aligns with the full model's loss geometry, yielding an emulator that faithfully replicates the training signals essential for adaptation. Across multimodal VQA, LLM personalization, and image generation, TALE consistently outperforms forward-only emulators and achieves performance competitive with full-model fine-tuning. Furthermore, it demonstrates robust generalization capabilities to unseen users and tasks, supporting decoupled server-side generation for on-device use.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.