acceptodds
Under review as a conference paper at ICLR 2027

EVESkillDebate: Evidence-Verifiable Multi-Agent Debate with Self-Evolving Agent Reasoning Skill

Abstract

Multi-Agent Debate (MAD) can improve language-model reasoning through iterative proposal, critique, and revision. However, agreement under a fixed protocol does not ensure that decisions are evidence-backed, nor does it convert recurring failures into persistent procedural improvements. To address these limitations, we propose a novel paradigm, EVESkillDebate, which consists of Evidence-Verifiable Multi-Agent Debate and Self-Evolving Agent Reasoning Skill. The debate module aims to construct candidate-specific reasoning paths that bind candidate-level claims to document spans and knowledge-graph edges. A frozen verifier assesses citation validity, evidential support, contradiction, and graph consistency, after which an evidence-aware judge aggregates the debated proposals and verification signals. The evolution module contrasts successful and failed (i.e., higher- and lower-utility) execution trajectories to identify recurring skill-level failures and propose targeted updates to the agents' external inference-time policies while keeping the underlying LLM frozen. Candidate skills are evaluated on subsequent batches and deployed only when they satisfy a predeclared feedback gate; post-deployment regressions trigger rollback. Across five public datasets spanning four domains and three LLM families, EVESkillDebate delivers its strongest and most consistent gains on evidence-grounded ranking tasks while achieving high evidence traceability. Evolved skills win 72.4% of non-tied comparisons and achieve 93.7% non-regression, showing that EVESkillDebate converts trajectory feedback into beneficial skill updates in the MAD systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.