acceptodds
Under review as a conference paper at ICLR 2027

Used Does Not Mean Needed: Self-Attention in Interleaved Flow-Matching Action Experts

Abstract

Does a performance drop after deleting self-attention establish that an action expert needs it? We study flow-matching vision-language-action models whose action experts interleave separate cross-attention (CA) and self-attention (SA) modules. With CA retained, we distinguish deleting SA from a fixed dense-adapted policy from adapting a policy with SA disabled throughout. In GR00T N1.7 on LIBERO-10, dense inference, post-hoc deletion, and No-SA adaptation achieve 93.27%, 85.33%, and 93.60% success. In SmolVLA at two integration steps, their four-suite means are 76.40%, 23.84%, and 88.38% under matched adaptation configurations that differ only in the SA switch. Interventions in the fixed dense policies further show that dependence is conditional: early SA changes from a negative to a positive observed contribution when late SA is retained. This characterizes learned reliance without identifying the mechanism of successful adaptation. No-SA deployment reduces measured latency by 16.97% and 23.40% under the two models' respective timing protocols. Together, these results show that post-hoc dependence does not establish downstream architectural necessity in the studied CA–SA action experts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.