acceptodds
Under review as a conference paper at ICLR 2027

Closed-Loop Agentic Reinforcement Learning for Vision-Language Driving Agents

Abstract

Vision-language driving policies commonly reason through a fixed chain of structured questions before acting, asking the same questions at every frame. On Bench2Drive-VL, a closed-loop benchmark that scores driving and reasoning on the same episodes, a model answering this chain drives worse than one answering nothing: which questions a frame warrants is itself a decision, and a fixed chain never makes it. Adaptive-reasoning methods do make it, but optimise the accuracy or cost of the answers where they are produced; the value of a question is realised only in the trajectory that follows. We formalise reasoning-chain selection as a closed-loop Markov decision process, which makes this mismatch explicit. On this basis we introduce DynaChain, an agentic driving policy that decides at every frame which questions to ask before committing to an action. We train it with closed-loop agentic reinforcement learning, under a hybrid advantage that combines the episode's driving outcome with a dense per-step reward derived from the benchmark's scoring rule. DynaChain raises the Driving Score of its base model from 58.93 to 79.15 while asking 2.52 questions per decision instead of seven, and outperforms two adaptive-reasoning methods. Our analysis shows that the gain comes from which questions the agent asks rather than how accurately it answers them, and that it requires both learned selection and closed-loop training. We will release our source code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.