Learning When to Listen: Closed-Loop LLM Advisory for Cooperative Multi-Agent Reinforcement Learning
Abstract
Large language models (LLMs) can serve as tactical advisors in cooperative multi-agent reinforcement learning (MARL), yet existing advisory channels fix the adoption decision outside the executing agents and provide no feedback on whether advice was followed or effective. We organize the interface around three dimensions: who decides adoption (D1), at what semantic granularity (D2), and what feedback the advisor receives (D3). We propose Closed-Loop Gated Advisory (CLOGA), a lightweight adapter for frozen recurrent executors that resolves all three: (i) semantic gating decomposes each directive into heads and lets each agent modulate adoption per dimension through learned, state-conditioned gates; (ii) a feedback channel returns per-head gate coefficients, relative reward, and situation drift to the frozen advisor for in-context revision without weight updates; and (iii) corrections are bounded residuals at the action readout rather than writes into recurrent memory. The system assumes a centralized advisory service at execution while preserving local recurrent backbones; this is an explicit channel assumption, not strictly communication-free decentralized execution. We establish conditional exact readout recovery and a readout-state deviation bound; both are pointwise results at a fixed history, not guarantees about joint-policy return. On SMAC and SMACv2 with MAPPO and HAPPO executors, CLOGA raises the weakest base by +0.25 and the imperfect-tier average by +0.09, while ungated injection of the same advice degrades three of four controlled maps; in a gate-type×feedback comparison, per-head gating converts feedback into a +0.024 gain against +0.006 for an otherwise identical learned scalar gate, suggesting that granularity and observability interact.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.