RUBATO: A Run-Ahead Asynchronous Compression Framework for Agentic RAG
Abstract
Agentic retrieval-augmented generation interleaves reasoning and retrieval, causing the model input to grow throughout a request. Prompt-level context compression limits this growth, yet existing agent pipelines place each compression task between two agent steps, blocking subsequent reasoning or retrieval. In our measurements, compression under synchronous execution accounts for 43%–54% of single-request service time. Longer service times increase worker occupancy and queueing under concurrent serving. We present Rubato, a run-ahead asynchronous framework that removes trajectory compression from the request's critical path. A context-construction rule (CCR) combines the latest committed compressed context with subsequent uncompressed steps, allowing the agent to advance while compression runs in the background. A prefix-monotone criterion (PMC) accepts a result only for an active request and only when it extends the prefix covered by the committed context, preserving consistent model inputs despite late or out-of-order completions. Across four compression methods, six question-answering datasets, and three agent backbones, Rubato reduces dataset-averaged mean single-request latency by – and median job completion time under concurrent serving by – while preserving average answer quality. These results support treating compression placement as a serving decision separate from the compression method. Our code are openly accessible to support future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.