Config2Weights: Generating LLM Initializations from Architecture Configurations
Abstract
Training large language models (LLMs) from scratch remains prohibitively expensive, motivating growing interest in weight-space generative modeling as a route to direct parameter synthesis. Existing generative approaches, however, are largely confined to small vision models or low-rank adapters, while heuristic model expansion methods rely on rigid, architecture-specific transformations. We propose **Config2Weights**, a generative initialization framework that learns from a collection of pretrained LLM weights to initialize new architectures conditioned solely on their architectural configuration. Config2Weights comprises two complementary variants: a direct configuration-to-weight mapping and a conditional latent flow operating in the latent space of a pretrained conditional variational autoencoder over model weights. We evaluate Config2Weights on both LoRA adapters and full-model weights across in-distribution and out-of-distribution settings, including unseen architectures and configurations. Across these settings, Config2Weights achieves competitive performance against existing weight-expansion methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.