When the Second GPU Helps: Topology-Aware Mode Selection for Narrow-Interconnect LLM Serving
Abstract
Modern LLM serving assumes NVLink-class interconnects, yet most multi-GPU deployments (consumer workstations, PCIe-only cloud VMs) link cards at a few GB/s. On such a narrow link, does a second GPU help, and in which mode? Replica places a full model copy per card, roughly doubling throughput when the context fits one card; staged split partitions the layers, extending reach but moving the per-token attention state (KV cache, growing with context) over the slow link every step. We show that this decision reduces to a calibrated discrete decision rule over two measurable coefficients—portable after per-engine recalibration: the per-token transport tax τ and the replica's fixed-overhead gap ΔA. Their ratio gives a crossover context ctx* = ΔA/τ; we are explicit about identifiability: the decision is robust to ±15% coefficient error (first flip ≈31%), while the point estimate of ctx* is the fit's least-identified parameter (LOCO 120–1,281). We measure both at four link settings across three hosts (6.65/13.3/17.6/25.4 GB/s), four engines, two models, up to k = 4 no-NVLink GPUs: the split tax is sub-linear (flat through k = 3), rejecting the shared-root (k−1)² prediction on distinct-bridge hosts, while replica scaling is near-linear (×3.89). Because τ is system-controllable, the coefficients also serve as an interface: a per-token cost reduction (an fp8-KV plugin, quality-failing in the tested configuration) measurably shifts ctx* rightward; we identify where the crossover breaks (fast links, ΔA < 0). The abstraction is embodied in an automatic decision layer that gates on memory, calibrates online, and drives llama.cpp within its 32,768-token admission window while reaching 65,536-token contexts on two 12 GB consumer GPUs through its own 4-bit staged path, with no manual quantization or RoPE tuning; the reach is not exclusive—manual vLLM and ExLlamaV3 configurations also admit 65K, and are faster there.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.