Monitoring Emergent Reward Hacking During Generation via Internal Activations
Abstract
Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of completed responses, it remains unclear whether such behavior can be identified during generation and whether activation-based monitors remain valid as models are further adapted. We propose an activation-based monitoring approach that detects reward-hacking signals from internal representations as a model generates its response. Our method trains sparse autoencoders on residual stream activations and applies lightweight linear classifiers to produce token-level estimates of reward-hacking activity. Across multiple model families and fine-tuning mixtures, we find that internal activation patterns reliably distinguish reward-hacking from benign behavior, generalize to held-out mixed-policy adapters, and exhibit model-dependent temporal structure during generation. We further show that monitor validity under model post-training updates is direction- and model-dependent: complete monitors can transfer from SFT to RL without requiring component replacement, while reverse RL-to-SFT transfer can benefit from updating downstream monitor components. Finally, we evaluate the frozen monitor on an external single-turn EvilGenie setting, comparing internal activation scores against an independent behavioral judge without retraining or recalibration. Together, these results suggest that activation-based monitoring can provide a complementary signal of emergent misalignment while remaining useful across adapter variation, post-training model updates, and external evaluation settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.