acceptodds
Under review as a conference paper at ICLR 2027

TACIT: Internalizing Collaboration via Training-Time Experience Transfer

Abstract

Multi-agent reasoning with large language models (LLMs) has unleashed strong potential for complex tasks. However, existing paradigms, such as multi-round debate, rely heavily on verbose and transient inference-time communication, incurring substantial token overhead and sometimes redundant interactions that undermine performance. Consequently, communication remains a recurring runtime patch rather than internalized competence among agents. To this end, we introduce **TACIT** (**T**raining-time multi-**A**gent **C**ollaboration via **I**mplicit experience **T**ransfer), a framework that enables LLMs to *learn from the exemplary when confronted with the flawed* through two-stage preference optimization. Specifically, LLM-based agents first develop *self-verification* by directly assessing their own answers during generation. This endogenous confidence signal then *gates implicit collaboration* in the second stage: uncertain agents acquire superior peer trajectories as confirmed experience, thereby constructing agent-specific preference data for co-training. At inference, TACIT requires *no explicit communication*, relying only on lightweight confidence-driven routing and answer aggregation. Experiments across diverse reasoning benchmarks show that TACIT elicits reliable and general self-verification for genuine collaboration, surpassing representative baselines in task performance. Instead of test-time indiscriminate dialogue, targeted experience transfer makes TACIT a truly *tacit* paradigm — *collaboration, once learned, need not be spoken*.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.