CLASHarness: Benchmarking Multi-Agent Conflict in Computer Use via a Controllable and Verifiable Generation Pipeline
Abstract
Vision-language models (VLMs) are increasingly deployed in multi-agent systems to automate complex computer-use workflows, prompting numerous benchmarks to evaluate and advance agent capabilities. However, existing benchmarks mostly assume consistent, cooperative environments. In contrast, real-world collaboration rarely runs on consistent information: peer or sub-agents frequently report outdated states, issue contradictory instructions, or compete over shared permissions and files, forcing an agent to check others' outputs and verify ground truth in the operating environment. Evaluating such conflict resolution remains challenging: manual task authoring cannot scale across complex systems, while generating tasks via unconstrained LLM generation produces narrative drift, answer leakage, and unsatisfiable grading rubrics that leave tasks unsolvable. We introduce CLASHarness (Conflict-Loaded Agent Scenario Harness), a controllable framework for generating verifiable multi-agent conflicts in computer use, accompanied by a benchmark of 600 multi-round tasks. By combining hierarchical scenario planning with an eight-step verified micro-pipeline, CLASHarness pre-commits decisive visual facts before rendering and compiles hybrid code and semantic graders directly from executed reference trajectories, guaranteeing that every task is solvable by construction. Evaluating 15 frontier models on CLASHarness reveals three core bottlenecks. First, when outputs of sub-agents conflict with the environment, weaker models blindly follow misleading information, whereas stronger models will inspect visual evidence to verify ground truth. Second, resolving conflicts is much harder than finding them: while agents can spot inconsistencies, actively fixing the system state drops task success significantly. Third, coordination protocols remain a universal bottleneck: even frontier models fail over 58% of closing actions, frequently forgetting to release shared locks or hand off tasks to sub-agents. CLASHarness provides a scalable, verifiable foundation for diagnosing coordination failures and building reliable multi-agent systems. Code and Dataset will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.