acceptodds
Under review as a conference paper at ICLR 2027

Developmental Cognitive Interpretability: Modelling Value Learning in LLMs

Abstract

Narrowly fine-tuning an LLM to express one value can generalise to others it was not directly trained on, making the full effects of complex alignment post-training methods hard to predict. Understanding these effects is becoming increasingly important to ensure correct value alignment of LLMs as they become more capable and are deployed into a wider variety of situations, even during training. To aid this understanding, we propose developmental cognitive interpretability (DCI), a framework that seeks to predict how training influences an agent's behaviour, without having to carry out that training. It models an agent's behaviour as arising from some latent state, according to a cognitive model, with a developmental rule determining how that state changes during training. We study this framework by training Qwen3.5-9B with DPO to express different behavioural values, and then evaluating the effect this has on its values more broadly. We find that changes in behavioural value expression can be compactly expressed: a cognitive model assigning a score to each value predicts held-out preferences, and changes in these scores are low-dimensional. Furthermore, a developmental cognitive model fitted to a population of agents each trained to express a single value can predict the expressed values of agents with more complex training pipelines. Overall, our results show that compact models of the latent dynamics of behavioural value generalisation in LLMs can provide accurate predictions, even out of distribution. Such models offer a way to investigate and potentially anticipate the broader behavioural effects of narrow alignment training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.