OmniTTT: Linear Complexity Multimodal Visual Generation with Test-Time Training
Abstract
Unified multimodal generative models jointly model text, images and videos within a shared context, enabling flexible visual generation but producing increasingly long token sequences dominated by visual tokens. Test-Time Training (TTT) offers a promising alternative to quadratic attention by compressing long contexts into an online-updated inner model with linear complexity. However, directly converting attention to TTT in pretrained generative models leads to substantial performance degradation under lightweight continual training. We attribute this difficulty to a modality asymmetry: language tokens are relatively few and inexpensive to retain with attention, yet difficult to compress without information loss, whereas visual tokens dominate computation and can be compressed effectively with the appropriate architectural design. Based on this observation, we introduce OmniTTT, a simple TTT-based architecture for efficient multimodal generation, which models visual interactions with TTT while retaining minimal attention for language. OmniTTT inherits weights from pretrained full-attention models and can be adapted through lightweight continual training without the need for training from scratch. Extensive experiments on text-to-image, text-to-video, and omnimodal reference-to-video generation demonstrate that OmniTTT substantially improves efficiency, achieving, for example, a 8.75 attention kernel and 2.87 full model speedup on reference-to-video generation with comparable quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.