acceptodds
Under review as a conference paper at ICLR 2027

SELA-TTT: Semantically Modulated Test-Time Training for Vision-Language-Action Models

Abstract

Context enables Vision-Language-Action (VLA) models to use past interactions and demonstrations during task execution. Test-Time Training (TTT) represents this information through recurrent fast-weight updates, but conventional KV-binding updates leave implicit how each interaction's semantics should guide its contribution to context. We introduce Semantically Modulated Equivalent Linear Attention for Test-Time Training (SELA-TTT), which uses explicit semantic signals to relate these contributions to task structure. Our method modulates the queries, keys, and values in TTT's equivalent, time-evolving linear-attention formulation, conditioning both context writes and reads. The resulting fast state carries semantic information forward to subsequent decisions. We instantiate this interface with subtask advancement, overall goal progress, and atomic-skill optimality, capturing complementary aspects of task progress and execution validity. Evaluations across simulation benchmarks and real-world task suites show improved success over TTT baselines, with equivalent linear-attention modulation outperforming injection into input features or raw TTT projections. Controlled label perturbations further support the lasting influence of semantics through their correspondence with interaction history. Beyond stronger task performance, SELA-TTT enables semantic-aware in-context learning from imperfect human demonstrations and generalization to previously unseen failures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.