acceptodds
Under review as a conference paper at ICLR 2027

Activation Diffusion, Revisited: Design Choices for Diffusion-Based Metamodels of LLM Activations

Abstract

Recently, diffusion-based metamodels of LLM residual-stream activations were proposed as priors for downstream tasks such as improved activation-based model control by projecting steered activations back to the data manifold, and unsupervised detection of human-interpretable concepts using the activation of the diffusion metamodel, similarly to sparse autoencoders. They also proposed a raw generative quality metric as the ability of the diffusion model to match the data mean and covariance, and showed initial evidence that all of these metrics tend to improve with scale. We take an in-depth look at design choices of training these diffusion models. In particular, (i) we find that informativeness of the chosen diffusion activation distribution matters greatly for the steering correction task, (ii) we propose a new diffusion parameterisation that results in particularly good generative quality scores on activation data, (iii) we provide a comprehensive analysis of the interplay between steering and diffusion projection strengths, and (iv)we further verify that the concept detection performance scales with increasing width and depth through an extensive hyperparameter sweep. Our results provide a basis towards practical real world usage of diffusion-based metamodels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.