SAPO:Sparse-Aware Prompt Optimization
Abstract
Large Language Models (LLMs) exhibit activation sparsity that varies with input prompt structure, creating an opportunity to jointly optimize prompts and model computation. We introduce Sparse-Aware Prompt Optimization (SAPO), a framework that combines prompt engineering with reduced-order computation to accelerate both homogeneous and heterogeneous inputs. This is important for physical and embodied AI, where prompts can combine natural language, numerical sensor data, system parameters, and equations. We show that explicitly organizing heterogeneous information induces contextual activation locality, producing more correlated and compressible activation representations across neighboring layers. SAPO exploits this structure through a transformation based on Koopman operator theory and the Mori–Zwanzig formalism, replacing -neuron continuous-depth layers with compact -dimensional Gated Recurrent Unit (GRU)-based blocks () and replacing iterative Ordinary Differential Equation (ODE) solves with parallelizable matrix operations. Experiments on large-scale foundation models show that reduced-order computation accelerates both prompt types, while prompt engineering amplifies its benefit for heterogeneous inputs. SAPO achieves training speedup for structured prompts compared with unstructured prompts, together with approximately lower aggregate test time while maintaining comparable task quality. These results demonstrate the synergy between prompt engineering, which exposes more compressible activation structure, and reduced-order computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.