acceptodds
Under review as a conference paper at ICLR 2027

From Sampling to Selection: Grounding LLM Reasoning with Predictive Models of Physical Systems

Abstract

Test-time sampling can place a correct answer in the candidate pool without providing a reliable way to identify it. Scientific workflows offer an additional source of evidence: predictive models of the physical system. We compare three placements of the same predictive-state evidence during LLM inference—post-generation candidate selection, one-step evidence-conditioned revision, and direct state-to-decision mapping when a task rule is specified—and ask whether such evidence helps and when it makes the language model unnecessary. In a controlled simulation-grounded study and an independently implemented diagnostic study on public combined-cycle power-plant data, candidate coverage far outruns majority voting: with eight matched samples a correct candidate is present for 91.3% of questions, while self-consistency returns 63.8% and predictive-state selection 81.3%, recovering 63.6% of the selection gap at no additional sampling cost. Two controls delimit the mechanism. Selection fails when the score evaluates the wrong task quantity and, more stringently, when the required state change is smaller than the predictor's own error. And when the task's answer is a deterministic function of the predicted state, applying the task rule directly to that state (95.0%) matches the best LLM interface (95.6%) with no language-model calls at all. The results give a conditional account of predictive-state grounding and a concrete criterion for when language inference adds value over the predictor alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.