acceptodds
Under review as a conference paper at ICLR 2027

Gacha Decoding: Eliciting Diverse Generations Through Instruction Following

Abstract

Language models often produce surprisingly homogeneous responses to open-ended prompts, with a model's diversity tending to decline as its capability increases. We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing approaches in diversity at equal quality (up to 2.3x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x), and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves–even as its entropy decreases and diversity under prior approaches declines. Together, our results highlight that instruction following, rather than token entropy alone, can serve as the engine driving response diversity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.