Remember or Recite? Pricing State Restoration for On-Device Multi-Agent LLM Inference
Abstract
Multi-agent large language model (LLM) workflows are moving onto personal devices, where several roles share one base model through low-rank adaptation (LoRA) adapters and take turns, each with its own key–value (KV) cache. Device memory holds only a few such sessions, so evicted state must be written to flash or recomputed later, a cost that dominates the workflow's latency and energy. Existing serving systems price this choice as data centers do, assuming symmetric memory and independent requests. Our measurements show that both assumptions fail on a device: flash writes over 130× slower than it reads, and the workflow graph already says which turns are ready, mergeable, and finished, so much state need never be written. We propose Recite, an on-device runtime that turns these two facts into three closed-form rules: a per-tier save-or-recompute threshold in which session length cancels, liveness read off the workflow graph, and batched decoding of ready same-role turns under a memory-traffic ceiling. From traces and device constants alone, the rules predict opposite decisions on the two flash tiers of one board and where batching hurts. On Jetson boards with two model families, Recite cuts debate and voting makespan 2.1–2.5× and energy 2.3–2.7× at equivalent accuracy against the strongest of five reference configurations, and runs up to 6.8× faster than the stateless ones across all workloads. Where both arms of a decision are measured, the rule picks the faster arm every time, and the cost model predicts the restore share of held-out configurations with R² = 0.99.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.