acceptodds
Under review as a conference paper at ICLR 2027

Gefen: Optimized Stochastic Optimizer

Abstract

AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook, thereby reducing AdamW’s memory footprint by up to 8× while maintaining the same performance, corresponding to a reduction of 6.5 GiB per billion parameters. Prior work reduces optimizer memory by sharing second moments across parameters grouped along the Hessian’s block-diagonal structure, but relies on hand-specified architectural rules and leaves unexplained why such grouping works. We prove that large mixed Hessian entries constrain the ratio of squared gradients toward one. This supplies the missing explanation, since a shared second moment is accurate when the squared gradients it pools are similar. It also shows the Hessian need not be computed: a block-structured Hessian forces the squared gradients to inherit that structure, so the blocks can be found in the squared gradients themselves. Gefen therefore infers block structure from the initial squared gradients, requiring no architecture-specific metadata or user-tuned hyperparameters beyond AdamW defaults. Gefen learns an exact histogram-based dynamic-programming quantization codebook and reuses the same blocks for first-moment scaling. Across diverse pretraining experiments, Gefen achieves the lowest peak optimizer memory among the compared methods that maintain AdamW-level performance. In single-machine and distributed training, the reduced footprint enables larger microbatches and substantially improves throughput over AdamW, making Gefen a drop-in replacement that can train larger models or use larger global batch sizes. We provide the complete Python implementation, including fused CUDA kernels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.