Under review as a conference paper at ICLR 2027
EFFICIENT GENERATIVE AI BY TUCKER NETS
Abstract
We propose a trainable token mixing transformation based on Tucker maps — which are of (quasi)linear complexity in the input. On experiments on ViT and GPTs we show that replacing attention layers by our Tucker maps caused % reduction in the number of FLOPs used in training — while almost never causing any deterioration of performance. We verify that the advantages of Tucker Nets on transformers, hold for both Adam(W) and Muon being used for training. Our Tucker layers are designed as drop-in replacements of the attention layers in any transformer architecture and thus pave a new way to efficient generative AI.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.