Toward a Science of Model Organisms
Abstract
Model organisms (MOs), models deliberately constructed to exhibit a behavior of interest, have become a central tool in alignment research. They allow us to study harmful behaviors not yet seen in deployed models and to develop and test evaluations, safeguards, and mitigations against them. This requires the MOs to have predictable properties. However, the field lacks a systematic science for building and evaluating MOs. Several methods exist to implant a behavior into a base model, but how these MOs behave differently remains unclear. We aim to bridge that gap by introducing MOBench, an automated evaluation suite testing MOs for properties (such as generality, robustness, coherence, etc.) identified via backchaining from how the MO is supposed to be used, as described by our framework. As a case study, we train MOs for the "impulsiveness" behavior, spanning interventions on the context, activations, and weights, and then evaluate them using MOBench. We find that Open character training (OCT) and fine-tuning work best for general behavior expression >95%, followed by DPO >60%. OCT is 1.3 x better than SFT for generality with lesser degradation in coherence, and prompting does not produce strong or robust MOs but has high introspection. Such trade-offs are invisible to a single behavioral metric, and we propose our property profile as a better basis for comparing model organisms. Together, these are a first step towards a deliberate science of model organisms. We open-source our code, evaluation suite, and organisms for the community to build upon.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.