T-WLA: Bridging Slow Vision and Fast Touch
Abstract
Contact-rich manipulation requires tactile corrections guided by visual information and past observations. We introduce , a world–language–action architecture that connects visual planning with tactile correction through cached visual context. At each visual update, the visual branch combines the current image with context from previous observations to generate an action chunk and update this context. Between visual updates, a lightweight tactile editor combines the cached context with recent tactile context to produce residual corrections without rerunning the visual branch. During training, future-image and future-pressure prediction help the model learn features useful for action generation and correction. We evaluate on tactile-conditioned hand-motion prediction using EgoTouch and Tachin, and on four real-world dexterous manipulation tasks. reduces finger-joint MPJPE by 30% relative to the strongest reported baseline and achieves an average success rate of approximately 50% across real-robot manipulation tasks. Videos and codes are available at https://anonymous-submission-20.github.io/T-WLA.github.io/ the project website.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.