Dispersive Sampling: Batch Diversity via Cross-Sequence Repulsion
Abstract
This paper introduces Fisher Diverse sampling, an inference-time approach to encourage the samples drawn from a language model to be semantically distinct from each other. To accomplish this, we use determinantal point processes (DPP) to couple autoregressive sampling across the batch dimension. DPPs are a standard method for defining probability distributions over sets of dissimilar elements, given pairwise similarity scores of those elements. In our setting, this requires us to specify similarity scores for each pair of possible output tokens. The challenge then is to find a numerical score that suitably captures the semantic similarity between outputs. We accomplish this by using the Fisher (information) geometry of the language model's output distribution to define a (context-sensitive) notion of similarity on the token unembeddings. Qualitative examples and empirical evaluation show that the method is highly effective at increasing the diversity of model outputs without degrading quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.