acceptodds
Under review as a conference paper at ICLR 2027

When Does Acting Improve Learning? Humans, Language Models, and the Role of Thinking

Abstract

Does learning from the same realized interaction depend on whether the learner chose the actions that produced it? We study this question in humans and language models using INTERGROUND, a matched-role paradigm in which an acting learner chooses consequential actions while an observer receives the same realized action–response record. Among 120 adults, acting improved transfer to novel and counterfactual situations by 5.8 percentage points (95% CI [2.6, 8.9]). Across 11 language model configurations spanning SmolLM2, Qwen3, Llama 3, and Gemma 3, however, standard online supervised adaptation produced only small acting–observer differences. In Qwen3, enabling thinking during learning increased the acting–observer gap across scales, with clear evidence at three of four tested scales; at 8B, the gap rose from 0.3 to 4.5 percentage points, a 4.2-point increase (95% CI [2.6, 6.4]). Randomizing whether the executed action matched the learner’s preferred action showed that thinking increased sensitivity to action–experience alignment. The same thinking-dependent pattern generalized to a structurally distinct causal-intervention environment at Qwen3-8B, where the gap increased from 0.4 to 2.9 points. Our findings show that matched trajectories do not guarantee matched learning: replaying an interaction alone does not always reproduce the benefit of generating it. For agent training, these results argue for retaining learner-specific context preceding each action, including reasoning and action preferences, alongside the observable trajectory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.