acceptodds
Under review as a conference paper at ICLR 2027

It Takes a Good Model to Train a Good Model: Generalized Gaussian Priors for Optimized LLMs

Abstract

While the rapid advancement of large language models (LLMs) drives increased scale, data transfer, and computational demands, the statistical structures of their weights, activations, and gradients—and their implications for initialization, training dynamics, and efficiency—remain largely unexplored. We empirically show that these quantities in LLMs are well modeled by generalized Gaussian (GG) distributions, and introduce a unified, end-to-end optimization framework grounded in this observation. Our contributions are threefold: (1) a GG-based initialization that aligns with trained model statistics, accelerating training and improving accuracy; (2) activation constrained training (ACT), which regularizes activation representations with GG constraints, reducing pipeline-parallel transmission overhead; and (3) gradient constrained training (GCT) that enforces low-bitrate gradient structures, enhancing communication efficiency in distributed training. Experiments across diverse architectures demonstrate consistently smaller, faster models with minimal communication overhead that match or surpass standard baselines. By anchoring LLM optimization in principled statistical modeling, this work advances efficient, scalable, and hardware-aware AI systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.