DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
Abstract
Many real-world tasks require coordination among multiple independent devices, but existing device-control benchmarks primarily focus on a single interactive environment. Large-scale executable benchmarks for evaluating such cross-device coordination remain scarce. We introduce DevicesWorld, a large-scale executable benchmark for cross-device coordination in heterogeneous environments, comprising 5,894 tasks across mobile, desktop, and IoT environments, including 5,614 cross-device tasks. DevicesWorld provides a unified environment for cross-device interaction and evaluation, enabling agents to execute tasks across multiple independent devices with automatic evaluation. We evaluate 13 agent baselines on DevicesWorld-Lite; the best achieves a Task Success Rate of 40.00%, well below the human reference of 84.00%. To further examine how these failures relate to single-device subtask execution capabilities, we select 60 cross-device tasks from DevicesWorld-Lite, decompose them into 237 single-device subtasks, and conduct independent tests using three representative models, comparing the results with end-to-end execution on the same original tasks. The three models achieve subtask success rates of 79.75%–90.30%; however, for each model, at least two-thirds of the original tasks whose subtasks all pass independent tests still fail end to end. Combined with trajectory analysis, these findings highlight task planning and decomposition, cross-device information linking and task-progress tracking, and feedback-driven replanning as directions for improving end-to-end cross-device execution. By releasing DevicesWorld, we aim to provide a unified infrastructure for future research on cross-device agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.