Acting for the Right Reasons: Decision-Cause Alignment in Trajectory Planners
Abstract
Vision–language models (VLMs) are increasingly adopted as end-to-end trajectory planners trained by behavior cloning, which aligns actions but not necessarily the reasons behind them and can therefore mask safety-critical failures. Existing methods infer a planner's reasons indirectly from representations, interventions, or behavioral associations. However, such evidence is not the decision-driving event itself; if the inferred reason changes with the diagnostic, it reflects the diagnostic–planner pair rather than a property of the planner. We instead define the *decision-driving event*, the world event that actually drives the planner's response, and call a planner **decision-cause aligned** when this event is the human-relevant event rather than a co-occurring cue. When the human-relevant event and the cue share an onset and response form, static behavior cannot separate their contributions. When the cue appears first, ordinary driving logs can identify its contribution under explicit response-latency and disturbance-isolation assumptions. Building on this result, we introduce **TREAT** (loca**T**e–cu**R**e–c**E**rtify for decision-c**A**use alignmen**T**): Locate measures temporal locking to candidate cues, Cure removes diagnosed dependence with audit-targeted decorrelation records, and Certify calibrates the audit by defect injection. Across nuScenes and Argoverse 2, four VLM families, the non-VLM planner UniAD, and 100+ trained instances, replica fleets show cue-ward dependence that is graded and population-structured. Although the confirmed case occurs in the most cue-ward fleet observed, it remains rare at the instance level; certificates therefore bind individual trained weights. **TREAT** isolates this naturally occurring defect where four standard diagnostics do not, repairs it with less task degradation than compute-matched continued training (ADE 2.72 vs. 3.22 m; 1.62 m before repair), and calibrates audit resolution with per-planner false-positive rates and detection limits at 80% power. Code: https://anonymous.4open.science/r/TREAT-5C14/README.md
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.