acceptodds
Under review as a conference paper at ICLR 2027

EBM-CoT: Energy-Based Calibration for Implicit Chain-of-Thought

Abstract

Large Language Models (LLMs) demonstrate strong reasoning capabilities through Chain-of-Thought (CoT) prompting, which enables step-by-step intermediate reasoning. However, explicit CoT relies on discrete token-level processes that are prone to error accumulation and often lead to rigid and inconsistent reasoning trajectories. Recent work has explored implicit reasoning in latent spaces, allowing models to perform internal multi-step reasoning before generating outputs, but such approaches lack explicit mechanisms to enforce reasoning consistency. To address this limitation, we propose , an Energy-Based Chain-of-Thought Calibration framework that refines latent reasoning trajectories using an energy-based model (EBM). By guiding latent representations toward low-energy regions associated with coherent reasoning paths, improves both reasoning accuracy and consistency without modifying the underlying language model. Experiments on mathematical, commonsense, and symbolic reasoning benchmarks demonstrate that EBM-CoT matches the performance of 10-chain Self-Consistency decoding using only a single reasoning chain, reducing the required sampling budget by 5x to 19x without modifying the underlying language model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.