Eliciting Model Behaviors by Optimizing the Latent Posterior
Abstract
In the latent posterior model of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations. We exploit this model in settings where it is exact, namely Bayes-filtered transformers (BFTs) meta-trained on sequences from a hierarchical prior, to introduce Posterior Prefix Tuning (PPT), a new method for eliciting behavior from a transformer: given a utility function on continuations, find a prompt under which the model's expected utility is large. For a BFT, the elicitation objective factors through the latent posterior, and we show that the gradient of a tilted surrogate is a covariance over the tilted posterior between the per-model utility and the score of the tilt. PPT optimizes the parameters of a distribution over hard prompts, drawing latent-model samples once from the BFT's prior via predictive Monte Carlo (PMC) and computing its gradient by importance sampling. The inner loop performs no transformer forwards and no backpropagation through the model, and the prior samples are utility-independent, so the cost of characterizing the BFT amortizes across utilities. We validate PPT and a Rao–Blackwellized variant on Beta–Bernoulli and reinforced urn BFTs across three utility families (reverse cross-entropy, frequency matching, Dyck validity).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.