memLLM-net: An Empirical Comparison of Wi-Fi, Shared-Memory, and RDMA Transports for On-Device LLM Context Delivery
Abstract
Beyond a single machine is where MemLLM’s [1] central argument stops holding: it argued that zero-copy shared memory eliminates the serialization overhead dominating on-device LLM inference, a claim scoped, by construction, to one machine. memLLM-net tests what happens at that scope’s edges. We build and validate, on real hardware: (i) an eRPC-style reliable transport for the case where the application and inference server are not co-located, exercised over synthetic loss and then a real home Wi-Fi link between two physical devices; (ii) the Linux memfd_create/eventfd shared-memory path MemLLM always described but never measured, completing its own stated evaluation gap; and (iii) a real RDMA transport via Soft-RoCE, requiring no specialized hardware. Five genuine defects surfaced and were fixed in building the Wi-Fi transport, including a retry-storm failure mode that reproduces, and complicates, the standard claim that RoCE livelocks on lossy links. Real shared memory beats the Wi-Ficapable transport by roughly 4× and real RDMA verbs beat our own Python shared-memory implementation by roughly 36× on the same machine — a measured ceiling this line of work falls well short of. We further argue, against an earlier framing of our own, that the KV cache should almost never cross the network at all: a thin client tracking only what changed, not the cache itself, is sufficient for the common architecture, which reframes what “KV-cache-over-Wi-Fi” should actually mean.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.