acceptodds
Under review as a conference paper at ICLR 2027

Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning

Abstract

When Large Language Models produce structured outputs such as travel plans, code solutions, or multi-step proofs, individual reasoning steps may appear correct while the output as a whole is globally invalid or non-optimal. Since valid candidates often exist in the candidate pool, reliable selection is a key bottleneck. We propose a decomposed energy function that combines a learned quality scorer with deterministic constraint penalties for verifying structured LLM outputs. The quality scorer is a heterogeneous ensemble of low-rank adapters on a single frozen encoder (3% trainable parameters); the ensemble mean ranks candidates, while its standard deviation estimates epistemic uncertainty. We also experiment with a two-pass strategy that uses uncertainty and constraint violations to trigger regeneration or abstention. Across five benchmarks (GSM8K, MuSR, TravelPlanner, TACO, and Knights & Knaves), our 149M-parameter verifier orchestrating a pool of 7-26B open generators reaches 67.7% on MuSR, compared with 68.0% for Claude Sonnet 4.6, and reduces TravelPlanner constraint violations by 53% relative to Opus 4.6. The results show that structural verification and pretraining-scale priors are complementary: deterministic verification is strongest when constraints are checkable, capturing signals frontier models cannot self-detect, while frontier priors remain stronger for narrative inference and code semantics. A cross-dataset confounding analysis confirms genuine quality discrimination on four tasks and identifies a model-identity shortcut on TACO, the code-generation benchmark. Last-layer retraining reduces the dominant generator’s selection share from 94% to 34% while retaining 88.6% pass@1 on TACO.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.