acceptodds
Under review as a conference paper at ICLR 2027

HADP: Game-Grounded Hierarchical Adaptive Dialogue Policy for Aligning Actions and Utterances in MOBA Team Coordination

Abstract

Competitive multiplayer games require AI teammates not only to choose useful actions but also to honor what they tell players. We study this say–do gap in MOBA team coordination and propose Hierarchical Adaptive Dialogue Policy (HADP). HADP grounds player requests in a shared Mission representation, generates coordinated action and dialogue responses, and reviews candidate outputs against evolving context and memory. We evaluate alignment at three levels: semantic consistency between utterances and issued plans, execution consistency between activated plans and agent behavior, and interaction consistency of commitments across turns. In the evaluated closed-loop scenarios in Honor of Kings, HADP achieves 91.5%, 86.3%, and 76.5% consistency at these levels, respectively, compared with 81.3%, 67.5%, and 51.5% for a flat policy with budget-matched self-refinement. Module ablations show that removing Understanding lowers semantic consistency by 18.0 percentage points, while removing Adaptive Review lowers interaction consistency by 15.3 points. These results support complementary roles for mission grounding and contextual review within the tested setting. Supporting trace-based experiments assess macro-action prediction and command quality after supervised fine-tuning, GRPO, and command-channel self-distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.