Encoding Prompts into Single Vectors via Linear Activation Aggregation
Abstract
Large language models process prompts by propagating a sequence of activations through dozens of layers before generating a response. How that sequence should be encoded into a single representation that keeps the relevant information of the prompt is largely unresolved. Existing approaches sit at two extremes: heavy learned compression methods for efficiency, and simple activation-engineering heuristics like the last-token activation, chosen for convenience rather than fidelity. Motivated by the advances in mechanistic interpretability, we show that a prompt can be represented as a linear combination of its own activations. We instantiate this approach with a lightweight MLP that predicts a weight for each token. This design encodes a prompt into a single activation vector whose information can be recovered by the frozen LLM when reinjected at an early layer. Despite its simplicity, the weighted-sum representation preserves prompt information more faithfully than end-to-end compression and provides an interpretable construction. Beyond its practical implications, our analysis reveals structure in the representation space of LLMs: (i) mid-layer representations transfer meaningfully to early layers, suggesting a degree of cross-layer compatibility in how information is encoded; (ii) a single activation vector encodes a quantifiable and recoverable amount of semantic information; (iii) a weighted sum of activations provides an effective and interpretable mechanism for single-vector prompt encoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.