LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
Abstract
Efficient language model deployment requires architectures that balance model quality with the energy and latency costs of the target hardware. We present LLMForge, a hardware-aware neural architecture search framework for pretrained language models. Its Infinite-Head Attention parameterization decouples query-head count from head dimensions and allows separate query/key and value widths. Together with variable key–value grouping and MLP widths, this defines a flexible search space for capacity allocation within and across layers. To evaluate candidates efficiently, LLMForge converts each pretrained model into an elastic supernet and jointly trains sampled subnetworks with shared weights, avoiding separate candidate training during search. Four interchangeable backends provide GPU measurements, accelerator modeling, joint model–chip optimization, and cost prediction from smartwatch measurements. Their cost evaluations guide multi-objective search alongside supernet validation loss, allowing the selected architectures to adapt to both the hardware backend and the optimization objectives. We evaluate the framework using five pretrained models spanning 135M-4B parameters. Across eight model–accelerator configurations, the search reduces modeled energy per token by a median of 8.4% relative to uniform scaling at matched supernet validation loss.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.