acceptodds
Under review as a conference paper at ICLR 2027

Thinking with Metacognition: Rebalanced Exploration-Exploitation in LLMs

Abstract

Despite recent advances in reasoning capabilities, Large Language Models (LLMs) often exhibit strikingly greedy behavior, exploring far less than humans in sequential decision-making tasks. We investigate this discrepancy across three bandit settings with human reference data. To characterize the underlying choice dynamics, we adapt a computational model of human exploration-exploitation (EE) behavior. Across ten LLMs, we find that their native reasoning (i.e., Thinking mode) amplifies sensitivity to estimated value, causing value differences to dominate choice while limiting the strength and persistence of uncertainty-directed exploration. Analyses of reasoning traces reveal that LLMs tend to compare value estimates and translate provisional judgments directly into choices, with little assessment of their reliability or evidential support. Such assessment and regulation of one’s own judgments constitute Metacognitive monitoring and control, whose limited engagement gives rise to a “value-to-action” shortcut. Motivated by this diagnosis, we scaffold metacognitive reasoning in LLMs by prompting them to assess value, uncertainty, confidence, and whether the current state calls for exploration or exploitation, and to use these assessments to regulate choice. Across LLMs and experimental settings, metacognitive reasoning rebalances sensitivity to value and uncertainty, promoting early uncertainty-directed exploration and improving subsequent exploitation. These gains persist in more challenging long-horizon four- and five-arm tasks, where improved optimal-arm selection and cumulative reward show that metacognitive reasoning can support adaptive exploration-exploitation over extended decision horizons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.