Thermodynamic Weight Decay: Global vs. Block-wise Control around Criticality
Abstract
Weight decay is typically applied using a fixed coefficient whose selection requires task-specific tuning. Inspired by dissipative structures in nonequilibrium thermodynamics, we propose Thermodynamic Weight Decay (TWD), a dynamic weight-decay method that views learning as a balance between energy influx and dissipation. TWD partitions a network into architectural parameter blocks, such as embedding, attention, and feed-forward modules, and treats the squared weight norm of each block as its internal energy. It sets a reference energy for each block according to its architecture and role, and drives the block energy toward a scaled target. On modular addition and multiplication tasks, TWD accelerates grokking by approximately 60-fold compared with training without control, while matching the performance of tuned static weight decay. Compared with global control, which constrains only the network-wide total energy, block-wise control prevents distortions in the energy allocation across blocks and accelerates grokking by approximately 20-fold. On Fashion-MNIST and CIFAR-10, TWD reduces test loss and the generalization gap while largely maintaining test accuracy. A broad sweep of the target-energy scale reveals four empirical regimes. Lowering the target initially accelerates grokking, but eventually weakens post-onset stability and robustness across random seeds; at still lower energy, grokking is substantially delayed or absent. Conversely, high energy leads to lazy feature learning, while the reference scale provides a balance between plasticity and stability. Experiments on GPT-2 (124M) show that block-wise control remains feasible at this scale, while also revealing that suitable energy targets depend on both parameter type and architecture. These results indicate that both the level and allocation of weight energy strongly affect plasticity and stability during learning, and may provide an operational indicator of criticality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.