SG-TTT: Scene-Grounded Supervision for Test-Time Training of World Action Models
Abstract
World action models (WAMs), which couple action generation with predictive world modeling, can degrade substantially under distribution shifts between training and deployment. Test-time training (TTT) offers a way to adapt these models using observations collected during deployment. For WAM adaptation under deployment shifts, future-prediction TTT typically matches each predicted future signal with its realized counterpart, without explicitly exploiting relationships among signals that describe the same realized future state, potentially limiting the informativeness of supervision. We formulate supervision for future-prediction TTT as a two-step process: organizing observed signals and their relationships into a structured supervisory reference, then translating this reference into adaptation objectives. We propose Scene-Grounded Test-Time Training (SG-TTT), a TTT method for WAMs with supervision grounded in scene geometry. SG-TTT first constructs a factual scene anchor from multi-view visual observations and the corresponding end-effector pose obtained after action execution, together with their spatial relationships. It then applies component-wise scene consistency to measure the scene-level spatial inconsistency introduced by each predicted future-state component and derives component-specific adaptation signals from the shared anchor. Extensive experiments on long-horizon LIBERO-10 suite under both standard and LIBERO-Plus distribution-shift settings, and real-world manipulation tasks show that SG-TTT achieves stronger overall performance than the frozen base WAM and representative future-prediction TTT baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.