TrustFlow: Counterfactual Information-flow Control for Reliable LLM-based Multi-Agent Collaboration
Abstract
LLM-based multi-agent systems improve complex reasoning by allowing multiple agents to exchange information and iteratively refine their solutions. However, unrestricted information sharing can also propagate unreliable reasoning and lead to error amplification and false consensus. Existing approaches largely focus on judging information reliability from current signals, such as agent confidence, consensus patterns, or external verification, but pay less attention to how information should be controlled to improve future collaboration outcomes. We propose TrustFlow, a counterfactual information-flow control framework for reliable LLM-based multi-agent collaboration. TrustFlow formulates collaboration control as a dynamic topology selection problem. First, it estimates a joint reliability belief over participating agents from observable interaction histories, providing a structured state beyond independent trust scores. Second, it evaluates alternative information-flow topologies through counterfactual rollouts and estimates their downstream effects. Third, it trains an outcome-aware controller with feedback-based policy optimization to select interventions, including preserving current reasoning, peer verification, and independent rethinking. Experiments across diverse reasoning benchmarks show that TrustFlow consistently improves reasoning performance and robustness while reducing inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.