What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives
Abstract
AI models serving a heterogeneous population need to be able to act according to the goals or principles appropriate to different users and contexts. Post-training has been shown to narrow the range of perspectives reflected in large language model outputs, but existing evidence primarily concerns behavior under neutral prompting. We show that post-training methods can additionally degrade models' ability to be steered in-context toward perspectives they were not trained to favor. In controlled experiments, we fine-tune policies to favor one side of cultural-value disagreements and evaluate checkpoints throughout optimization. The selected target becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing viewpoint declines. These findings reveal a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. We then develop and theoretically characterize a principled alternative that optimizes reward subject to a prescribed distribution over expressed perspectives and present stance-distribution matching as one practical implementation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.