KARMA: Epistemic Principles for Trustworthy Multi-Agent Collaboration
Abstract
Multi-agent LLM systems amplify collective capability but also amplify errors: without epistemic discipline, fabricated information propagates through agent networks, and agents cannot reliably distinguish credible from unreliable peers. From structural constraints on decentralized evidence processing, a constraint-to-failure-to-principle derivation yields KARMA, implementation-agnostic epistemic principles that govern both what an agent may assert from its own evidence and how it weights peers by track record. We then test whether this specification can be operationalized across two mechanistically distinct substrates. In LLM experiments on 150 multi-layer adversarial scenarios ( runs, 36,000 total judgments across six conditions), KARMA produces large improvements over the no-principle baseline for mid-tier models: to pp (Cohen's –, all ). Without principles, models perform near chance (49.7–53.7%); with KARMA, they approach frontier performance (90–99%). The frontier model achieves 100% under both conditions—the specification adds nothing when intrinsic capability suffices. Generic prompting strategies—caution, chain-of-verification, reputation, role authority—do not substitute for the specification: all fall significantly short of KARMA for non-frontier models (, –). The reputation heuristic is the strongest competitor but collapses on recency-conflict traps where, for two of the three mid-tier models, it falls below the no-principle baseline, precisely where temporal discounting—prioritizing recent outcomes over stale reputation—is required. A companion Bayesian simulation characterizes convergence and scaling behavior under the same specification across 7,000+ runs. The two mechanistically unrelated substrates—natural-language reasoning and an explicit Bayesian reputation mechanism—exhibit structurally aligned behavior (signal extraction, unreliable-source resistance, and capability-dependent activation) while sharing only the KARMA specification; both implementations were constructed by the same authors and therefore demonstrate realizability rather than independent convergence. Frontier models from multiple labs reach KARMA-consistent judgments without explicit principles, while capacity-constrained models incur a prompting-substrate floor—bounding useful operationalization to an intermediate capability regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.