acceptodds
Under review as a conference paper at ICLR 2027

ADE: Accurate, Inference-efficient, and Tunable Delta Compression for Task-specific Fine-tuned Foundation Models

Abstract

Given multiple task-specific fine-tuned models with the same base, how can we deploy them compactly while preserving each model's performance? Delta compression stores only the difference between a fine-tuned model and its base, enabling compact deployment of many models from a shared base. However, existing methods largely treat the delta as an isolated artifact: they ignore the base model, disregard inference computation, and produce a static result that cannot be further adapted. We propose ADE (Accurate Delta compression with Efficient inference), an accurate, inference-efficient, and tunable delta compression framework with (1) base-awareness, allocating parameters by the sensitivity of each module's output, including the base weights, for accuracy, (2) computation-awareness, aligning quantization with the direction of matrix multiplication for fast low-bit inference, and (3) task-awareness, supporting near-zero-overhead tuning on the compressed delta for low-cost adaptation. Under the same memory budget, ADE is up to 3.9% relatively more accurate than the best prior method, and has up to 18.2% lower inference latency than the fastest prior method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.