acceptodds
Under review as a conference paper at ICLR 2027

Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation

Abstract

Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is hindered by the computational complexity and global reduction bottleneck imposed by layer normalization. Recent methods attempt to bypass this by replacing normalization layers with hardware-friendly scalar approximations. However, these homogeneous replacements do not optimally fit to all layers' behaviour and rely on expensive model retraining. In this work, we propose a highly efficient, hardware-aware framework that utilizes genetic programming (GP) to evolve heterogeneous, layer-specific scalar functions directly from pre-trained weights. Coupled with a novel post-training re-alignment strategy, our approach eliminates the need to retrain models from scratch entirely. Our evolved expressions accurately approximate the target normalization behaviours, capturing 90–93% of the variance () compared to only 70–76% for homogeneous baselines, allowing our modified architecture to recover 84.32% Top-1 ImageNet-1K accuracy on ViT-B and 85.74% on ViT-L in only 20 epochs. By retaining near-baseline accuracy while eliminating the global reduction bottleneck, our approach achieves a strict reduction in both arithmetic complexity and off-chip memory traffic compared to standard LayerNorm, removing a primary barrier to the efficient deployment of ViTs on edge accelerators.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.