acceptodds
Under review as a conference paper at ICLR 2027

Chain-of-Attention: Field-Verifiable Reasoning for Traffic Vision–Language Models

Abstract

Traffic vision–language models can produce plausible answers without revealing when an event occurs, which actors are involved, or what visual evidence supports the explanation. We introduce Chain-of-Attention (COA), a phase-indexed record that binds five traffic phases to selected frames, pedestrian and vehicle boxes, and an event-level account. These explicit fields support separate training, scoring, reward optimization, and stress testing. On 48 fixed-camera WTS validation scenes, supervised training raises pedestrian IoU from .014 to .406 and vehicle IoU from .035 to .570 under the field-level evaluation with uniform sampling and no phase labels. Reward optimization preserves this grounding and improves blinded ratings of visual factuality, causal support, and evidence consistency. A phase-free control shows that supervision balance strongly affects reliable structured output. Balancing full-record and auxiliary examples raises strict schema completion from 3/48 to 15/48; COA reaches 16/48 under its corresponding strict schema. This experiment identifies task mixture as a major source of format failure, while COA provides an explicit interface for training and auditing evidence-bound predictions. We also expose a temporal shortcut: equal-per-phase sampling makes phase identity recoverable from input rank. Replacing images across 47 paired scenes collapses actor IoU while leaving the temporal score essentially unchanged, showing that temporal structure alone does not establish visual grounding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.