Designing Better Transformers: Using Amorphous Network Proxies to Engineer the Unified Highway Transformer
Abstract
Standard transformer architectures define information flow through rigid, serialized sequences of discrete operators. In deep or very deep implementations, this leads to PreNorm Dilution, or Latent Amnesia, in which deep layers become blind to early layer signals in the residual stream, stalling convergence. Existing solutions to PreNorm Dilution are computationally expensive or add engineering complications. This work utilizes unconstrained Amorphous Neural Network (AmNN) clouds as diagnostic proxies allowing analysis of information flow that emerges free from the constraints of discrete operators. These paths empirically confirmed the demand for historical latent context in deeper layers. Modeling this flow back into a discrete operator transformer resulted in the creation of the Unified Highway Transformer (UHT). The UHT introduces an parallel memory bus () that builds a history of latent context from the residual stream and makes that history available to deep layers. To counter initialization instability during training, a Dynamic Variance Annealer initially scales residual projections entering the bus, phasing out scaling after initial epochs. The UHT was evaluated against a constrained 120-layer, 9.7-million total trainable parameter architectural stress test () on WikiText-103. A gradient tracker revealed dynamic, high-variance, gradient norm activation at deep levels in the UHT, illustrating that the UHT restored the representational capacity of the deep layers. In contrast, a 120-layer Pre-LN control showed gradient stagnation in deep layers. The UHT reduced terminal perplexity in the tested configuration (118 for the Pre-LN to 66 for the UHT). An ablation UHT removing all layer normalization operations achieved a similar perplexity of 68. A 180-layer test of the UHT with layer normalization operations showed similar performance. The UHT provides a scalable, parameter-efficient architecture for extreme-depth modeling that solves gradient stagnation in deep networks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.