acceptodds
Under review as a conference paper at ICLR 2027

Latent Test-Time Computation with Energy Transformer Components

Abstract

Looped transformer blocks pass hidden states through shared weights multiple times, offering benefits in model size reduction and latent reasoning. While most work in this area has focused on traditional transformers, a related line of work on energy-based modeling introduced Energy GPT (EGPT) blocks, whose forward pass is explicitly the gradient of an energy, or negative log-likelihood, function. Energy-based models are therefore naturally looped, since inference corresponds to gradient-based optimization of the energy. This raises a natural question: can energy transformers be combined with traditional transformers, and does their gradient structure provide practical benefits? We find that the answer is yes. Hybrid EGPT-GPT models are substantially more robust to reductions in loop steps than purely GPT-based looped models, leading to better test-time compute scaling. We identify two qualitatively different computation modes. Variants with energy attention exhibit distributed refinement, in which predictions improve progressively at each iteration, with correct-token rank improving from 110 to 3 across blocks. In contrast, some GPT variants exhibit concentrated construction, in which useful predictions are deferred until the final iterations, with correct-token rank remaining above 250 until the last block. The practical consequence is that distributed refinement enables compute savings at inference time: a simple convergence-rate criterion saves 19% of measured FLOPs with no quality loss for energy-attention models, while concentrated-construction models cannot be halted early without collapsing. These findings replicate across scales from 200M to 1B parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.