acceptodds
Under review as a conference paper at ICLR 2027

Self-Consistency via Marginal Sharpening

Abstract

Test-time compute has become a standard way to improve language-model reasoning without changing model weights, by allocating additional computation to sampling, search, or refinement before committing to an answer. A recent line of inference-time power-sampling methods pushes this idea by amplifying high-probability model outputs, typically doing so by sharpening complete generated sequences. We argue that this is the wrong object for reasoning models: a complete output entangles a latent reasoning trace with a final answer, even though many distinct traces may support the same answer. We introduce *Marginal Sharpening*, which instead sharpens the answer marginal induced by aggregating over latent reasoning traces. This gives an answer-level analogue of self-consistency, but as an inference target rather than a post-hoc voting rule. For integer sharpening strengths, the target has an exact multi-trace decomposition. We use this identity to derive a simple parallel auto-regressive approximation that samples multiple traces and generates one answer whose tokens are jointly supported across them, avoiding both complete-sequence search and exact string aggregation. Across math and code benchmarks, Marginal Sharpening improves over standard sampling and is particularly strong for code generation, where semantically equivalent solutions often differ in surface form. Compared with power sampling, Marginal Sharpening delivers stronger performance with substantially lower runtime, supporting answer-level sharpening as a more effective test-time objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.