acceptodds
Under review as a conference paper at ICLR 2027

HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Agency

Abstract

LLM agents increasingly work with humans who act as collaborators, not passive task providers. We introduce *HAS-Framework*, a graph-based framework for Human-Agent Systems (HAS) that represents humans and LLM agents as first-class participants with explicit roles, permissions, communication paths, and action authority. We then build *HAS-Bench*, a benchmark of 397 tasks across six domains in which the human agency level, interaction channels (clarification, feedback, and control), and user persona are configurable. *HAS-Bench* scores task outcomes and the collaboration process, including clarification quality, feedback utilization, control-request justification, safety, initiative, and interaction cost. Across five LLMs with simulated users, Equal Partnership (A3) scores 8.4 points higher in Pass@1 and 11.5 points higher in Task Score than Full Automation (A1) on average. In GPT-4.1 ablations, each channel is strongest on a different problem pattern: clarification when information is hidden, feedback when outputs need revision, and control when actions need authorization, where control alone reaches a 100% Safety Rate. Bargaining Pass@1 varies by up to 63.0 points across user personas. Existing benchmarks treat the user as a fixed task provider or study a single user in a fixed role; *HAS-Bench* instead makes the degree and form of human participation a controlled experimental variable. Code is available at: https://anonymous.4open.science/r/HAS-Bench-958B/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.