Frozen Text Embeddings with a GNN is “All You Need” for Text Classification
Abstract
Recent work in text classification constructs token-level graphs from the embeddings of a frozen language model and trains a graph neural network (GNN) on these graphs to perform downstream classification. We show that existing approaches lack consistent performance across standard benchmarks, with no single graph construction strategy dominating across datasets. To address this limitation, we introduce HILT-GCN, a simple hierarchical token-level graph architecture that models multi-layer interactions via layer-specific virtual hubs at linear complexity - HILT-GCN achieves the strongest and most stable overall performance among frozen token-level approaches across five single-label benchmarks. Compared with strong fully fine-tuned baselines based on both small and large text encoders, HILT-GCN achieves competitive results on four out of five datasets, falling behind only on Ohsumed. By using frozen embeddings, HILT-GCN supports any new text backbone without encoder fine-tuning, achieving high accuracy while substantially reducing the parameter and training overhead of large-scale models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.