How to Train Your Energy-Based Transformer: Understanding Stability in EBMs
Abstract
Diffusion and autoregressive models generate images, video, and text at scale. Energy-based models (EBMs) cast generation as energy minimization. They learn a scalar energy that scores how well a candidate output matches the data, and energy-based transformers (EBTs) scale this approach by implementing the energy with a transformer. An EBT generates by refining a prediction through a sequence of gradient descent updates on the energy, starting from a noisy or corrupted input. Training differentiates through this refinement sequence, the same computation as backpropagation through time (BPTT) in recurrent networks, and it is often unstable. We analyze this instability. Each backward step multiplies the gradient by a factor set by the step size and by the curvature of the energy with respect to the prediction. We bound the resulting growth and convert the bound into a curvature score computed during training. We compare full BPTT with two lower-cost gradient estimators, truncated BPTT, which differentiates only the last few updates, and Anticipated Reweighted Truncated Backpropagation (ARTBP), which cuts the sequence at random points and reweights the surviving terms to remove the bias introduced by the cut. In image denoising with three model sizes, we find that the best estimator changes with scale. ARTBP gives the highest final training reconstruction quality with the small model, while truncation to the last three updates is best for the base and large models. With the objective and forward trajectory fixed, we find that the truncated gradient changes from closely aligned with the full gradient early in training to pointing in the opposite direction later. The curvature score separates failed from successful completed runs and predicts training stability over a horizon that shrinks with model size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.