acceptodds
Under review as a conference paper at ICLR 2027

Remember or Recite? Pricing State Restoration for On-Device Multi-Agent LLM Inference

Abstract

Multi-agent large language model (LLM) workflows are moving onto personal devices, where several roles share one base model through low-rank adaptation (LoRA) adapters and take turns, each with its own key–value (KV) cache. Device memory holds only a few such sessions, so evicted state must be written to flash or recomputed later, a cost that dominates the workflow's latency and energy. Existing serving systems price this choice as data centers do, assuming symmetric memory and independent requests. Our measurements show that both assumptions fail on a device: flash writes over 130× slower than it reads, and the workflow graph already says which turns are ready, mergeable, and finished, so much state need never be written. We propose Recite, an on-device runtime that turns these two facts into three closed-form rules: a per-tier save-or-recompute threshold in which session length cancels, liveness read off the workflow graph, and batched decoding of ready same-role turns under a memory-traffic ceiling. From traces and device constants alone, the rules predict opposite decisions on the two flash tiers of one board and where batching hurts. On Jetson boards with two model families, Recite cuts debate and voting makespan 2.1–2.5× and energy 2.3–2.7× at equivalent accuracy against the strongest of five reference configurations, and runs up to 6.8× faster than the stateless ones across all workloads. Where both arms of a decision are measured, the rule picks the faster arm every time, and the cost model predicts the restore share of held-out configurations with R² = 0.99.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.