Simple Video Context Scaling of Robotic Policies
Abstract
Intelligent robotic agents should exhibit strong memory capabilities over very long temporal horizons. Memory can be facilitated by directly feeding video history as context to the policy, but doing so comes with a significant challenge. Video memory spanning more than even just a few seconds quickly becomes computationally difficult: simple token math mandates tokens of context just for minute of fps video memory with current standard methods. To address this, in this work we show that by adopting an early-fusion-like model architecture, even very long videos can fit completely and token-efficiently in transformer contexts, changing the token math substantially. On simulated memory benchmark tasks our method, Scaled Robotic Memory via Early-Fusion (SRoME), greatly surpasses the token efficiency of history representations compared to baselines, while also being more performant. We also show that train step times and inference times are quick, with very favorable scaling behavior when memory context length scales. In short by tackling the main challenge that plagues video memory policies today, video context length, SRoME makes long video memory policy training easy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.