acceptodds
Under review as a conference paper at ICLR 2027

Delta Activations: A Representation for Finetuned Large Language Models

Abstract

The success of powerful open source Large Language Models (LLMs) has enabled the community to create a vast collection of post-trained models adapted to specific tasks and domains. However, navigating and understanding these models remains challenging due to inconsistent metadata and unstructured repositories. We introduce Delta Activations, a method to represent finetuned models as vector embeddings by measuring shifts in their internal activations relative to a base model. Clustering analysis shows that Delta Activations achieve strong separation of finetuned domains, significantly outperforming baselines such as flattened weights, salient parameter masks, and output embeddings, while being more lightweight and computationally efficient. Delta Activations also demonstrate desirable properties: it is robust across finetuning settings and remains useful even in dense local perturbation banks where many specialists lie close to the same pretrained weights. The method extends naturally to task embedding via few-shot finetuning for reliable model retrieval, and to model selection for merging by quantifying similarity. Delta Activations provide a practical foundation for organizing and reusing the growing ecosystem of post-trained models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.