acceptodds
Under review as a conference paper at ICLR 2027

Utility-Guided Residual Multi-Agent Reinforcement Learning for Closed-Loop Interactive Driving

Abstract

Closed-loop multi-agent driving without strong map structure (e.g., no lane markings or discrete lane IDs, and no fixed leader–follower roles) remains challenging for both classical motion planners and learning-based policies. Geometric interaction methods can under-specify objectives such as progress and comfort, while from-scratch Reinforcement Learning (RL) policies must discover useful behavior over large action spaces from limited safety-critical interaction data. We propose utility-guided residual multi-agent RL, where a discrete control scorer acts as an interpretable prior and a shared residual policy re-ranks that prior online. Every executed command remains an over the same discrete bicycle-kinematic candidate set, preserving the structured decision space and kinematic feasibility of the prior. The prior is a multi-term utility over one-step rollouts. We calibrate the prior independently on three naturalistic geometries (e.g., a weak-lane-discipline freeway, a no-lane-discipline urban roundabout, and an urban grid network) and evaluate closed-loop transfer across all three. Training follows a CaRL-style progress reward with multiplicative soft constraints and terminal contact; learned checkpoints use a common safety-first, then NAVSIM-style PDM selection rule, while deterministic inference applies a PDM-Closed accept-if-better gate. We benchmark against classical motion planners, geometric interaction methods, offline traffic-generation approaches, and matched online RL baselines under a common closed-loop evaluation protocol, and ablate residual parameterizations, the accept-if-better gate, forward collision lookahead, and traffic density. Under a matched low-budget training protocol on the freeway benchmark, the proposed residual achieves a higher PDM score and fewer collision events per episode than MAPPO and matched direct discrete RL. Ablations suggest that learned residual adaptation, not the accept-if-better gate, drives the gains; candidate-level residuals remain robust without collision lookahead, while the parameterized residual is stronger under dense interaction. These results support residual adaptation of a calibrated discrete utility prior as a sample-efficient approach to closed-loop interaction while retaining interpretable, kinematically feasible decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.