Transformers from Compressed Represenatations
Abstract
Compressed file formats are the cornerstone of efficient data storage and transmission, yet their potential for representation learning remains largely under-explored. We introduce TEMPEST (TransformErs froM comPressed rEpreSenTations), a method that exploits the structured byte-stream of compressed files to design an effective tokenization and encoding strategy. By leveraging this com- pact encoding, a vanilla transformer can directly learn semantic representations from compressed data streams. Our proposal reduces the number of tokens required for semantic classification, thereby lowering both computational complexity and memory usage. Moreover, we discover unique properties of compressed byte stream for representations learning, specifically, the existence of complementary information across compression levels. Through extensive experiments across diverse datasets, coding schemes, and modalities, we show that TEMPEST achieves accuracy competitive with raw data baselines while delivering efficiency gains in memory and compute.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.