MILO: Mode-wise Likelihood Optimization for Post-Training Diversity
Abstract
People often ask language models for multiple candidate solutions, but sampling more responses does not necessarily produce more ways of solving a problem: many samples may repeat the same underlying strategy. We introduce Mode-wise Likelihood Optimization (MILO), a post-training method for increasing the diversity of correct solution strategies while preserving accuracy. MILO groups verified-correct responses that use the same strategy into correct modes, gives each mode its own log-likelihood term, and combines the resulting mode-wise objective with pooled correctness using a weight . Its mode-wise update gives each observed mode equal total weight, rather than favoring modes simply because they appear more often. We show that this update directly targets correct-mode coverage. For finite rollout groups, its expected mode-wise update is the gradient of a -weighted sum of expected CorrectModes@, where CorrectModes@ counts the distinct correct modes among samples. At a local correctness optimum, a small MILO step with improves this coverage objective to first order when its gradient is nonzero, while correctness changes only at second order. On Qwen3-4B-Base GSM8K, improves both single-sample accuracy and CorrectModes@, raising coverage from to . Increasing to raises coverage to while keeping single-sample accuracy comparable to the base model and pass@ nearly unchanged. The gains extend across datasets, sampling budgets, and independent mode labelers, and are also reflected by labeler-free diversity metrics. A controlled verifier-blind-spot experiment further illustrates how broader correct-mode coverage can help candidate pools retain usable correct answers when a verifier fails on common strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.