acceptodds
Under review as a conference paper at ICLR 2027

Causal Concept Gating for End-to-End Driving: Faithful Concept Actuation Is Not Effective Concept Intervention, and Why

Abstract

What does a named concept channel buy a competitive closed-loop driving planner, how much of it is free, and what does intervening on it do? We answer with a hybrid latent–concept planner and measurements on it. A multi-modal encoder predicts 25 named scene factors, a prior-initialized, end-to-end-learned graph maps them to 8 named action gates, and the gates bias each metric head of an anchor scorer that also reads the scene latent directly, so one can trace a factor's effect through a gate to a metric, and the channel is an auditable control path rather than a bottleneck. The planner is comparable to the state of the art on navtest (0.8985 EPDMS-v2; 0.9030 with an inference-time collision rerank independent of the channel), and a three-seed, same-budget ablation cannot distinguish the graph from a black-box gate MLP, so the findings concern a competitive policy. Three findings follow. The channel has a free-amplitude boundary. Its amplitude is an inference-time scalar, and a pre-registered rule raises it fourfold, to a tenth of the selection signal (half on time-to-collision), at no measurable cost, with the first measurable cost at sixfold. It is a faithful actuator. do-interventions on the gates move the emitted trajectory as the gate names prescribe at every amplitude (longitudinal Spearman ρ = 1.00, lateral CIs excluding zero), factor interventions shift the anchor distribution in the prior's direction across 11 checkpoints (r = 0.78) and 21–57× more strongly on factors active in the scene, and a permutation null with a second, non-additive read-out separates learned from architecturally forced alignment. Faithful actuation is not effective intervention, and this dissociation is our main finding. Under the standard test-time-intervention protocol, replacing all 25 factors by ground truth changes the score by less than 0.001, and a matched control shows that the gate edge a factor intervention pushes sets its effect, not whether the assertion is true. We locate the loss at the read-out, which maps both the additive coupling's scene-independent push and a multiplicative (FiLM) coupling's scene-specific perturbation to a single direction and flattens them. What the channel does support is auditing, reading back a missed factor on a real logged at-fault-collision scene, separating a mislabelled concept from a model error, and repairing a blind safety-factor detector at no detectable planning-score cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.