acceptodds
Under review as a conference paper at ICLR 2027

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

Abstract

Model organisms (MOs) — language models trained to exhibit undesired or unnatural behaviours — are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT). Prior research has shown that interpretability methods can easily identify hidden behaviours in these MOs. However, recent work suggests that such post-hoc training methods may make interpretability unrealistically easy. We investigate this claim by constructing a suite of 61 - and -based MOs trained with seven different techniques, including standard post-hoc SFT, post-hoc DPO, and more realistic integration of MO data into the OLMo post-training DPO phase. We also use these to train 96 further variants via knowledge distillation. We use these MO variants to benchmark activation oracles, activation steering, the logit lens, and the Jacobian lens. Our findings show that (i) MO interpretability depends strongly on training objective, target behaviour, model architecture, and training data generation pipeline; (ii) substantial variance remains even after controlling for differences in the strength of target behaviour expression; and (iii) both our more realistic and knowledge distillation often yield less interpretable MOs than standard post-hoc methods. Our results cast substantial doubt on the validity of current MOs as interpretability proxies and suggest actionable steps towards more robust benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.