MatchCredit: State-Matched Credit Assignment for Adaptive Multi-Agent Reasoning
Abstract
Multi-agent reasoning (MAR) can improve reasoning reliability through additional collaboration, but adaptive routing requires determining which collaboration operator is worthwhile at each intermediate state. Existing query-level labels or trajectory-level returns provide only coarse supervision for this decision, as the effect of an operator is confounded by the reached interaction state and subsequent routing choices. We propose MatchCredit, a state-matched credit assignment framework for efficient adaptive MAR. MatchCredit expands alternative operators from the same reached state and evaluates their downstream continuations under a shared prefix, thereby reducing prefix-state confounding. Terminal correctness and additional suffix cost are then converted into state-local pairwise operator preferences to train a lightweight router. Collaboration trees are used only for offline supervision construction, while test-time reasoning follows a single adaptive trajectory. Experiments on six reasoning benchmarks show that MatchCredit substantially reduces inference cost while preserving or modestly improving reasoning accuracy, yielding a favorable accuracy–cost trade-off over representative MAR methods. Averaged across the six benchmarks, MatchCredit achieves 86.66% accuracy with 2,193 tokens per query.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.