Activation Diffusion, Revisited: Design Choices for Diffusion-Based Metamodels of LLM Activations
Abstract
Recently, diffusion-based metamodels of LLM residual-stream activations were proposed as priors for downstream tasks such as improved activation-based model control by projecting steered activations back to the data manifold, and unsupervised detection of human-interpretable concepts using the activation of the diffusion metamodel, similarly to sparse autoencoders. They also proposed a raw generative quality metric as the ability of the diffusion model to match the data mean and covariance, and showed initial evidence that all of these metrics tend to improve with scale. We take an in-depth look at design choices of training these diffusion models. In particular, (i) we find that informativeness of the chosen diffusion activation distribution matters greatly for the steering correction task, (ii) we propose a new diffusion parameterisation that results in particularly good generative quality scores on activation data, (iii) we provide a comprehensive analysis of the interplay between steering and diffusion projection strengths, and (iv)we further verify that the concept detection performance scales with increasing width and depth through an extensive hyperparameter sweep. Our results provide a basis towards practical real world usage of diffusion-based metamodels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.