FitBench: Benchmarking Human–AI Collaboration across Option-Picking-Driven Decision Styles
Abstract
Existing benchmarks primarily measure Large Language Models (LLM) or incorporate Harness's task capabilities and performance, which can accurately evaluate the agent's ability and efficiency in reasoning and thinking tasks. However, humans in reality have different decision-making personalities and preferences. Using a single evaluation criterion to assess AI cannot truly reflect the value of this performance in real-world applications, nor can it be used to measure human-AI collaboration efficiency. It is jointly determined by the task itself, AI performance, and human decision-making personality. Therefore, if it is to be evaluated, improvements need to be made to existing benchmarks by incorporating human decision-making personality variables. To this end, we propose **FitBench**, which represents human side differences as six structured Human slots under the same task conditions, and systematically evaluates human-AI collaboration performance under different decision-making personalities, while comparing the impact of different LLM and Harness configurations. FitBench models the collaboration process as three interconnected subsystems and introduces an optional Interactive Structured Assertion (iSA) interaction adapter to enhance the observability of slot differences. In the current experimental results, the range and variance of success rates across the six slots for the same actual task show significant differences; both increase in the recorded iSA control. In addition, we can select the optimal Harness configuration for each human decision-making personality slot while keeping the LLM layer fixed, thereby significantly improving the efficiency of AI collaboration with humans of different decision-making personalities. Thus, FitBench provides a unified evaluation framework for separating model performance from human decision-making differences, analyzing the conditional interaction between Human Slot and AI execution configuration, and further studying collaboration configurations tailored to different decision-making personalities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.