The Pro-Worker AI Benchmark: Measuring Whether LLM Assistance Preserves the User's Cognitive Role
Abstract
Capability benchmarks do not record whether a model's interaction pattern keeps the user's cognitive role. We introduce the Pro-Worker AI Benchmark (PWB), which draws on the human-AI complementarity literature to define eleven dimensions of interaction behavior, scored on 320 prompts (200 single-turn probes in thirteen professional domains, 16 five-turn scenarios of 80 turns and 40 adversarial stress tests) by a three-judge LLM panel. Four dimensions fall below quadratic-weighted among the judges and are excluded from the composite Pro-Worker Index (PWI); three blind human annotators also agree far less among themselves on those four (Krippendorff against ), and their consensus agrees with the panel median at . Across seven open-weight models from six families the default is answer-first: PWI ranges from 20.4 to 37.8 out of 100, baseline cognitive forcing is zero for five of the seven, and GPT-6 Sol and Gemini 3.1 Pro share the default. A system prompt that describes the scored behaviors raises the index to 50.9-82.8 ( = +25.2 to +53.0; paired over the composite's 120 prompts, ), which measures instructability: on three models a generic prompt of matched length, which instructs answer-first, adds +3.2 points, while a prompt written from a lay brief recovers 39% of the gain and a checklist of the rubric's surface features 68%. A responder that only asks questions beats every baseline, so we add PWI-D, which credits asking first only when a conversation with a simulated user delivers the requested output by the second turn: on six models PWI-D keeps 87-96% of the gain, and on three it drops the question-only responder to 3-8 points. The effect persists in five-turn scenarios ( = 1.61), under adversarial pressure where declining is defensible ( = 1.54), and under a judge from outside the evaluated families. The checks also find costs: on requests the prompt itself says to answer directly, prompted models ask first 47% of the time, and the judged calibration dimensions reward hedging language while the confidence models state on exact-answer questions does not move. PWB measures interaction form and does not establish effects on worker skill or decisions; we specify the study that would. All prompts, rubrics, judge logs and code are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.