Atxserve: Towards Transaction-Aware Inference Serving for Multiagent Systems
Abstract
Language model agents increasingly collaborate in shared workspaces to solve complex tasks. However, concurrent updates of multi-agents can compromise consistency or invalidate costly inference. A naive approach is to model these operations as database transactions and apply conventional concurrency control. But this direct transfer is poorly matched to agent workloads, since prolonged model inference dominates the interval between read and write operations. Optimistic validation detects conflicts only after expensive inference has been performed, while pessimistic locking can suppress useful concurrency throughout inference. This mismatch creates a fundamental tension between preserving consistency and GPU efficiency. In this paper, we address this mismatch by shifting the optimization focus from state access to model execution. Specifically, we reformulate agent concurrency control as a GPU scheduling problem and introduce Atxserve. Atxserve constructs a conflict graph before model inference, orders conflicting requests, and admits compatible requests together. Admitted requests then bind current snapshots and retain protection through validation and commit. This design avoids inference that would be invalidated by conflicts, and preserves concurrency among non-conflict transactions. To evaluate Atxserve across diverse concurrency scenarios, we develop RMWBench, which emulates multi-agent workloads with controllable conflicts. On workloads based on real-world repositories, aggregate comparisons show that Atxserve improves useful throughput over optimistic concurrency control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.