acceptodds
Under review as a conference paper at ICLR 2027

When Does a Neural Sampler Actually Need the Target Score? Isolating the Role of Langevin Parameterisation

Abstract

Nearly every effective neural sampler for unnormalised targets injects the target score into its drift, a design known as Langevin parameterisation. he2025notrick showed that removing it causes severe mode collapse and concluded that an explicit score term is required; their comparison, however, changes the score term and the velocity parameterisation at once. We re-run it inside a single architecture family, varying only the score's role. The term is indeed causal—coverage drops from modes to —but it works just as well as an ordinary input feature, and the capacity does not substitute for it, so what matters is the availability of the gradient rather than the residual structure around it. Most unexpectedly, within this setting the score is needed only during a specific interval of the sampling trajectory: restricting it to is numerically identical to having no score at all, recovers full coverage, and the last alone is worthless again. We give a partial account. The optimal control provably converges to the target score as , which explains the design but deepens the puzzle; a transport bound then shows the cost of redistributing mass inside a window of width diverges as , providing a control-cost mechanism consistent with the failure of late windows, and varying the mode separation moves the threshold consistently with the implied . The failure of early windows remains unexplained, and we report it as open rather than attach one of the two intuitive stories our own experiments falsify.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.