CIPI: When to Integrate Peer Feedback in Multi-Agent LLM Reasoning
Abstract
Conventional round-based multi-agent LLM systems integrate peer feedback at predefined interaction rounds, without considering whether the receiving agent needs an immediate correction or whether the feedback can wait. This fixed timing may delay necessary corrections or interrupt reasoning that is already correcting itself. We introduce CIPI (Conditional Integration of Peer Information), a protocol that lets the receiving agent decide when to use feedback after it arrives. CIPI stores peer feedback and uses the agent's current reasoning state to decide when to review it. The agent reviews feedback early when a contradiction or misunderstanding of the question could affect the next reasoning step and is not already being corrected. Otherwise, feedback is saved for later review. The model checks proposed corrections before accepting them and records accepted changes to guide later reasoning. Early corrections do not prevent further review, and a final check takes place before the answer is generated. The key idea is to use peer feedback when the receiving agent needs it. We evaluate CIPI on MuSiQue, HotpotQA, MATH-500, OlymMATH, APPS, and HumanEval, covering multi-hop question answering, mathematical reasoning, and code generation. We compare CIPI with output-level message passing, adapted Exchange-of-Thought Report, and message passing augmented with Self-Refine, evaluating task performance and measured token usage. Under the evaluated role-specific thinking settings, CIPI achieves the highest reported scores across the six benchmarks, improving over message passing with Self-Refine by 0.84–9.00 percentage points, generally with higher token usage. These findings suggest the value of selectively integrating peer feedback during ongoing reasoning to support timely correction without unnecessary disruption.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.