Process Matters more than Output for Distinguishing Humans from Machines
Abstract
Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online settings. Existing approaches largely evaluate whether a system can produce behavior or responses indistinguishable from those of a human. This approach follows the focus on the *output* of a machine as a criterion for determining whether it can think, as suggested by Alan Turing. Cognitive science provides an alternative approach: considering the *process* by which that behavior is produced. To evaluate whether differences in cognitive mechanisms can reliably distinguish humans from machines, we introduce a process-based framework, the *Process Turing Test*, and evaluate it across a battery of cognitive tasks spanning decision-making, working memory, and planning. These tasks, such as *mental rotation* and *sequence prediction*, yield process-level measures that reveal how behavior is generated, complementing conventional measures of overall task performance. We also include multiple CAPTCHA tasks in the battery to compare the experimental paradigms to challenges deployed in the real-world. Across the battery, process-level features provide substantially stronger discriminative signal than performance metrics alone, reliably distinguishing humans from agents even when task performance is matched (mean process-feature classifier AUC = 0.88). To assess agentic process limitations, we conduct a controlled red-teaming study comparing off-the-shelf frontier agents (Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro), Centaur (a large language model fine-tuned on 10.7M human decisions), and two task-specific fine-tuning methods applied to models of different sizes and families: **action-level supervised fine-tuning (A-SFT)** and **process-level fine-tuning (P-SFT)**, which directly optimizes process features. We find that broad fine-tuning on human choices makes task processes more human-like relative to off-the-shelf frontier agents, and task-specific process-level fine-tuning further improves human-like behavioral mimicry, though this advantage largely disappears under cross-task transfer when process targets do not naturally generalize across tasks. These results highlight process specification as a central bottleneck in achieving human-like cognitive processes in machines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.