acceptodds
Under review as a conference paper at ICLR 2027

LifelongCUAWorld: Benchmarking Computer-Use Agents for Lifelong Learning in Hybrid Environments with CLI Knowledge

Abstract

Computer-use agents have advanced in task execution, but reliable lifelong assistance remains underexplored. We introduce **LifelongCUAWorld** to evaluate the capability profile of *lifelong computer-use agents* through 180 expert-level tasks across 19 applications, with GUI/CLI/code interfaces and a structured CLI knowledge base. *Expert* evaluates independent tasks through verifiable deliverables; *Profile* tests acquisition and reuse of hidden preferences across task sequences for the same user; *Revision* provides staged feedback on agents' drafts and evaluates feedback requirements and preservation of previously correct content. Ten evaluated systems show partial professional competence, but none strictly completes an Expert task. DeepSeek Harness raises DeepSeek-V4.1-Flash's Expert score from 0.098 to 0.239. Six systems never request clarification, and preference-related delivery compliance reaches at most 6.3%. Agents can complete feedback rounds and preserve correct content, yet feedback compliance reaches at most 17.2%. Ablations support the CLI knowledge base and benchmark design: task-relevant CLI knowledge improves Expert scores, supplied prior user information improves Profile compliance, and Revision ablations support assessing feedback compliance and content preservation separately. Current systems still fall short of reliable lifelong assistance; LifelongCUAWorld provides a systematic evaluation platform for advancing *lifelong computer-use agents*. [Anonymous Project](https://anonymous.4open.science/r/LifelongCUAWorld-anonymous-38A2)

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.