Logic-Native Transformer: Mapping Transformer Computations to Trainable Logic Gate Networks
Abstract
The growing demand for deploying Transformers on resource-constrained edge devices motivates the design of hardware-efficient neural architectures. Differentiable logic gate networks offer one such approach by directly learning compositions of the Boolean operations underlying digital hardware, enabling representation learning in feedforward, convolutional, and recurrent architectures. Extending this approach to Transformers requires computing attention with Boolean operations and overcoming the optimization difficulties of deeply stacked logic gates. To address these challenges, we propose a Boolean attention mechanism based on XNOR similarity, top-k selection, and majority vote, and introduce XOR residual connections. Integrating these components with learnable logic gate modules yields a scalable logic gate network vision transformer (LGN-ViT) with a unified Boolean inference architecture. On CIFAR-10, LGN-ViT reaches a depth of 96 learnable logic gate layers, 6.4 times that of the convolutional logic gate baseline. For width scaling, our analysis shows 26.8% higher retention of marginal accuracy gains for LGN-ViT than for this baseline under a 3.35-fold increase in trainable parameter count.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.