Train for Many, Update with One: Improving Multi-Sample Reasoning via Fenchel-Bregman Posterior Learning
Abstract
Multi-sample reasoning measures *candidate coverage*, the probability that at least one of candidates is correct. However, minimizing the average loss of individual attempts does not directly optimize candidate coverage. We therefore formulate *finite- posterior learning*, which trains a distribution over solver parameters to reduce the probability that all independently generated candidates fail. To optimize this objective efficiently, we propose a Fenchel-Bregman (FB) approach that requires only one posterior draw per update, rather than the used by direct Monte Carlo optimization. FB first uses an exact Fenchel reformulation to separate problem-specific weights from the posterior update, enabling one-draw posterior updates. It then applies a closed-form Bregman update to track these weights through an exponential running average of observed failures, reusing the same draw. We instantiate FB with IVON, a variational optimizer for Gaussian posteriors over solver parameters, yielding **Fenchel-Bregman IVON (FBI)**. Across diverse recursive and autoregressive reasoning tasks, FBI improves candidate coverage. In a matched comparison, it also achieves comparable Pass@10 while using only 6.8% of the direct Monte Carlo baseline's measured training time per update.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.