acceptodds
Under review as a conference paper at ICLR 2027

Train for Many, Update with One: Improving Multi-Sample Reasoning via Fenchel-Bregman Posterior Learning

Abstract

Multi-sample reasoning measures *candidate coverage*, the probability that at least one of candidates is correct. However, minimizing the average loss of individual attempts does not directly optimize candidate coverage. We therefore formulate *finite- posterior learning*, which trains a distribution over solver parameters to reduce the probability that all independently generated candidates fail. To optimize this objective efficiently, we propose a Fenchel-Bregman (FB) approach that requires only one posterior draw per update, rather than the used by direct Monte Carlo optimization. FB first uses an exact Fenchel reformulation to separate problem-specific weights from the posterior update, enabling one-draw posterior updates. It then applies a closed-form Bregman update to track these weights through an exponential running average of observed failures, reusing the same draw. We instantiate FB with IVON, a variational optimizer for Gaussian posteriors over solver parameters, yielding **Fenchel-Bregman IVON (FBI)**. Across diverse recursive and autoregressive reasoning tasks, FBI improves candidate coverage. In a matched comparison, it also achieves comparable Pass@10 while using only 6.8% of the direct Monte Carlo baseline's measured training time per update.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.