SAE-STM: A Sparse Autoencoder-Augmented Structural Topic Model for Interpretable Pairwise Preference Prediction
Abstract
Reward models trained on pairwise human judgments drive RLHF and DPO, but they reveal little about what in a response makes it preferred. Neural topic models offer a compact, readable vocabulary but are too coarse to capture fine-grained preference factors; sparse autoencoders (SAEs) recover fine-grained features but are fit post hoc, with no link to the prediction task or to any structured summary. We introduce SAE-STM, which couples the two: a Top- SAE over the embedding difference of a response pair produces a sparse code that, following the Structural Topic Model, shifts the prior of a ProdLDA topic model, and a classifier predicts the human preference from the resulting topic proportions alone. Across six pairwise-preference datasets and four topic granularities, topic proportions and SAE features are complementary: a probe on both outperforms a probe on either alone. Finally, we check that the explanations are usable by someone outside the model: an LLM judge that sees only the two responses and the validated natural-language traits of their dominant topic recovers the human label with 53–74% accuracy, above the 50% chance level in all 24 dataset–granularity settings. SAE-STM thus provides preference predictions that come with a validated topic-and-feature explanation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.