PROFF: Prompt-Adaptive Reduced-Order Feed-Forward Computation for LLM Agents
Abstract
Modern LLM agents repeatedly process long prompts containing system instructions, tool schemas, and interaction history while generating short outputs. This makes prefill a major inference bottleneck, with feed-forward (FF) blocks accounting for a large fraction of its compute. Existing FF acceleration methods largely rely on neuron-level sparsity, which is less effective for modern gated activations such as SwiGLU and GeGLU. We present PROFF (Prompt-adaptive Reduced-Order Feed-Forward computation), which replaces the standard FF block with a reduced-order FF block. PROFF compresses the hidden state into a low-dimensional space, performs mixing in that space, and projects it back to the original hidden dimension. A lightweight predictor selects the reduced dimension for each agent call, adapting FF compute to the input. The surrogate is trained offline through per-layer distillation, while the remaining model parameters are frozen. Across four LLMs and multiple agentic benchmarks, PROFF retains near-baseline task quality while improving prefill throughput by 2.15–2.23 across prompt lengths and model architectures. Across multi-call agent workloads, PROFF further achieves 1.98–2.31 lower end-to-end task latency. These results show that prompt-conditioned reduced-order FF computation provides a practical approach to accelerating prefill-heavy agentic inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.