MaRePID: Multi-Agent LLM Reasoning via Distributed PID Regulation
Abstract
Multi-agent large language model (LLM) systems increasingly solve complex tasks through multi-round reasoning, where corrections from one agent can influence the subsequent reasoning of others. Such evolving interactions make it challenging to determine how feedback should guide corrections throughout the reasoning process. We introduce **MaRePID**, an inference-time framework that regulates multi-agent LLM reasoning using distributed proportional–integral–derivative (PID) control. For each agent, MaRePID maintains a bounded control state updated from evaluator-derived feedback and neighboring control summaries, and maps the resulting state to textual directives that steer subsequent reasoning. By jointly accounting for the current mismatch, its accumulated history, and its recent change, MaRePID dynamically adjusts prompt guidance while capturing interactions among agents. We further introduce prediction-triggered sensing to determine when new evaluator feedback is necessary, reducing unnecessary evaluator queries. Across MultiAgentBench, AgentWebBench, MATH, and BigCodeBench, the full-sensing, always-update variant of MaRePID consistently outperforms the evaluated fixed and adaptive baselines, achieving an improvement of up to 18.3 points over the base multi-agent system on AgentWebBench. These results demonstrate that MaRePID enables effective inference-time regulation of multi-agent reasoning without modifying model parameters or the underlying communication protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.