MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Abstract
Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between evaluation and deployment where personal assistants are expected to work across a user’s whole digital life, including their context, historical data, and logged-in accounts. This gap is widest on web tasks, where live web evaluations cannot exercise sites that require logging in or personal information, the kind of site a real personal assistant has to drive. We introduce MYPCBENCH, which tests computer-use agents as personal assistants on a Linux desktop populated with 17 simulated real-world web applications and a full desktop stack, all seeded for one canonical persona, Michael Scott from The Office. We define 184 tasks in this environment, each inspired by a real request drawn from the OpenClaw community, and benchmark ten closed- and open-weight models, each given the same computer and shell tools. The highest perfect rate is 58.2% (Claude Opus 4.6), with GPT-5.6-sol at 55.4% and a slightly higher rubric score. We use LLM-as-a-Judge to verify tasks, where our rubric judge agrees with human grading on 94.5% of sampled rubric items. Model failures cluster on tasks that span many applications and on long trajectories, where personalization stresses an assistant the most. We release the environment, task set, agent harness, and judge anonymously at https://anonymous.4open.science/r/MyPCBench-21FF/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.