Steer, Then Classify: Adaptive Document-specific Representation Steering for Text Classification
Abstract
Text classification is a fundamental natural language processing (NLP) problem, spanning many tasks such as intent and topic classification, toxicity and stance detection, entailment, factuality assessment and many more. Despite substantial advances in large language models (LLMs), modern text classifiers still largely follow the same paradigm, that encode the input with a pretrained language model and train a classification head on top of the resulting representation. This leaves the internal representations of the model largely passive during inference. In this paper, we introduce a different and novel approach to text classification based on representation steering. Rather than directly classifying a fixed document representation, our method dynamically modifies the representation of each input document individually along learned steering directions to make task-relevant information more separable in representation space. Building on this idea, we propose AdaSteer, an adaptive steering framework that determines how and where to intervene in an LLM's representations for each input. Across the Taarof and Multilingual Toxicity Benchmark datasets, AdaSteer consistently improves classification performance over conventional classification heads and existing steering baselines. Notably, when both approaches are initialized from the same underlying LLM, AdaSteer achieves a 1.89% average improvement in F1 score over the standard classifier baseline across four datasets (i.e. on the Irony, Taarof, Multilingual Toxicity and Sentiment datasets), while requiring comparable trainable parameters. Our experiments further show that per-document adaptive steering produces more discriminative representations and that the gains persist across model scales and datasets, providing a complementary alternative to conventional discriminative fine-tuning, and establishing representation steering as a promising paradigm for text classification. Our entire codebase will be released for facilitating reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.