acceptodds
Under review as a conference paper at ICLR 2027

Video as Experience: Watch, Act, and Rewatch for Computer-Use Agents

Abstract

Multimodal large language models have enabled increasingly capable computer-use agents that perceive graphical user interfaces, reason over instructions, and execute actions in interactive environments. However, existing agents largely operate on information available within the current interaction, while procedural knowledge from demonstrations is typically incorporated through separately extracted action trajectories. We introduce VEX (Video as EXperience), a framework that treats video as a unified representation of experience for both learning from external demonstrations and reflecting on an agent’s own behavior. VEX follows a simple watch-act-rewatch paradigm: an agent extracts task-relevant guidance from tutorial videos, acts in the target environment, and subsequently verifies its progress by rewatching its execution as video. We instantiate VEX with video-SALMONN-VEX, an 8B-scale audiovisual computer-use agent built on video-SALMONN 2+, and introduce Action-Grounded Guide Optimization (AGGO), a reinforcement learning method that trains the model to generate procedural guidance according to its utility for downstream action prediction. We further augment OpenCUA, OSWorld, and CUA-Gym with more than 24K real-world tutorial videos to support learning and evaluation from video experience. Experiments on OSWorld and a complementary subset of CUA-Gym show that VEX consistently improves computer-use performance, with video-SALMONN-VEX achieving a 15% relative improvement in task success rate over its backbone on OSWorld. Moreover, guidance generated by video-SALMONN-VEX improves Gemini-3.5-flash by 10% relatively on OSWorld. These results demonstrate the potential of video as a common experiential interface for acquiring procedural knowledge and verifying interactive behavior in computer-use agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.