TIDE: Selective State Exposure for Low-Budget Robot Transfer from Passive Video
Abstract
A low-budget manipulation controller benefits from interaction phase, while direct coordinates for the camera motion and edit boundary that produced a video observation need not be action inputs. Transfer Interface with Differentiated Exposure (TIDE) makes this distinction architectural: action-free pretraining learns persistent slot variables, vector-quantized temporal state, egomotion, and a cut-reset gate, while a frozen export map admits only slots and temporal tokens to the grounded policy. A same-backbone intervention tests this partition by fitting a separate action head for each visible coordinate set. With 100 grounded CALVIN episodes and 100 episodes per RLBench single-task head, heads that jointly receive egomotion and cut coordinates attain success lower by 0.119 and 0.125; withholding the temporal token costs 0.180 and 0.158. A matched-GPU-budget retrain without the two acquisition branches also attains success lower by 0.085 and 0.103. TIDE reaches strict five-instruction-chain success on CALVIN and mean success on RLBench, exceeding a compute-, width-, and head-matched action-free V-JEPA transfer by and . The advantage persists with one shared RLBench policy and across 8 ManiSkill2 tasks; Voltron requires as much grounded interaction to reach the same two-suite target. These results establish policy visibility as a separate design axis from predictive representation in the 25–100-episode regime.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.