Decomposing and Steering Functional Metacognition in Large Language Models
Abstract
Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, but whether this reflects a behavioral artifact or deeper internal structure remains unclear. We propose that LLMs maintain a decomposable space of *functional metacognitive states*: internal variables encoding evaluation awareness, self-assessed capability, perceived risk, computational effort, audience expertise, and intentionality. Across multiple reasoning models, these states are linearly decodable from residual-stream activations with accuracy scaling in model capacity. Template-generalization tests reveal that single-framing probes often read wording rather than state, whereas template-diverse probes transfer zero-shot to paraphrased and fully implicit framings for five of six dimensions. Probe-direction steering further shows that causal use is dimension-selective: suppressing intentionality and self-assessed capability degrades accuracy dose-dependently, with the intentionality effect surviving length control, whereas audience expertise is reliably decodable yet causally inert within tight equivalence bounds. Decodability, behavioral use, and generality are therefore distinct properties. Benchmark performance reflects not only task competence but also the model's metacognitive configuration; our framework provides a mechanistic basis for studying and controlling these states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.