MOVI-Agent: Learning Queryable Software Models from Instructional Videos for Professional GUI Automation
Abstract
Vision-language model (VLM) based agents can operate everyday applications through their graphical user interfaces (GUIs), yet they break down on professional software such as Darktable and KiCad, where dense panels, strict operational ordering, sparse on-screen guidance, and heavy domain terminology demand knowledge that pre-training does not supply. We present MOVI-Agent, a training-free agent that learns how a program works before operating it. From instructional videos it compiles a persistent, queryable software model: a function tree of what the software can do, a state graph of how its interface states connect, and an action database of how each demonstrated step is performed. Execution then becomes structured navigation over this model, and a semantic state backtracking mechanism rewinds to the last verified state whenever it drifts off track. We further construct ProbeGUI, a diagnostic benchmark for professional desktop software: 86 tasks in ten applications with step- and subtask-level error attribution, multi-path matching that credits alternative valid routes, and quantified recovery from injected errors. On ProbeGUI, no baseline exceeds task success; MOVI-Agent more than triples the task success of its GPT-5.4 backend ( to ), surpasses the strongest baseline on all four metrics with significant step- and subtask-level margins, and recovers from of injected errors. A matched control that receives the same videos as flat text shows where structure matters: it leaves step accuracy within 2.3 points but raises task success from to and error recovery from to .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.