AccessWorld: A Benchmark for Computer Use Agents Operating Assistive Technologies
Abstract
Computer use agents typically perceive graphical interfaces through screenshots or accessibility trees and act through mouse clicks and keystrokes. Many people with disabilities instead operate the same interfaces through assistive technologies (ATs), such as screen readers, voice control, and screen magnification. As a result, the action and observation spaces available to a user can differ greatly across interaction modalities. Assistive agents that can operate these technologies could collaborate with the people who rely on them, audit whether software actually works with ATs, and guide users through interfaces that are hard to navigate. There is also no reason to assume a GUI built for sighted mouse users is the right interface for an agent. Yet there are currently no tools or benchmarks for building and evaluating agents that operate ATs. We present AccessWorld, a benchmark of computer use tasks on macOS that agents complete by using system native assistive technology. Agents operate on the environment through an OpenClaw harness exposing each AT as an agent tool interface. We evaluate five models across five interaction modalities and find substantial differences across ATs in both task success and interaction cost. Through VoiceOver, macOS's screen reader, one model nearly matches its standard computer-use success, but requires substantially more actions and cumulative input. At magnification, three of five models exceed their standard computer-use accuracy, while increasing magnification to reduces accuracy for all five models. Voice Control produces the lowest accuracy for every model, with the largest drops on multi-item, spreadsheet, and text-editing tasks. AccessWorld measures how well agents handle the constrained observation and action spaces that AT users work with every day.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.