acceptodds
Under review as a conference paper at ICLR 2027

Decomposing and Steering Functional Metacognition in Large Language Models

Abstract

Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, but whether this reflects a behavioral artifact or deeper internal structure remains unclear. We propose that LLMs maintain a decomposable space of *functional metacognitive states*: internal variables encoding evaluation awareness, self-assessed capability, perceived risk, computational effort, audience expertise, and intentionality. Across multiple reasoning models, these states are linearly decodable from residual-stream activations with accuracy scaling in model capacity. Template-generalization tests reveal that single-framing probes often read wording rather than state, whereas template-diverse probes transfer zero-shot to paraphrased and fully implicit framings for five of six dimensions. Probe-direction steering further shows that causal use is dimension-selective: suppressing intentionality and self-assessed capability degrades accuracy dose-dependently, with the intentionality effect surviving length control, whereas audience expertise is reliably decodable yet causally inert within tight equivalence bounds. Decodability, behavioral use, and generality are therefore distinct properties. Benchmark performance reflects not only task competence but also the model's metacognitive configuration; our framework provides a mechanistic basis for studying and controlling these states.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.