Introspection or Post-Hoc Rationalization? A Cross-Model and Cross-Domain Study of LLM Metacognition
Abstract
Recent literature has increasingly studied the metacognitive and introspective capabilities of Large Language Models (LLMs), suggesting they may possess access to their internal states. However, a self-report can be *accurate* without being introspective. An accurate report may be caused by truly accessing the internal state (*introspection*), by the model recomputing that state in a new context (*self-simulation*), or by exploiting information that merely correlates with its internals (*post-hoc rationalization*), a route also known in LLMs as *causal bypassing*. In this paper, we empirically evaluate current introspection benchmarks under the core thesis that accuracy is *not* access. Because these evaluations score self-reports by accuracy, which alone cannot distinguish these routes, we introduce two complementary methodologies and apply them to test instruction-tuned LLMs from families, spanning B to B parameters. First, a *cross-model test* asks whether the rationalization route is available by giving external reporters exactly the inputs the self-reporter receives. We found that external reporters match self-accuracy, indicating the task is solvable via rationalization. A small residual self-advantage remains in some cases, which is compatible with a privileged route but cannot distinguish introspection from self-simulation. Second, an *introspection generalization framework* asks whether training on these evaluations builds a general route to internal states or only task-specific heuristics. Across task families and datasets, in-domain fine-tuning gains largely fail to transfer across tasks. Together, these results do not claim models cannot introspect, rather, they show that current evaluations cannot tell the routes apart, and that training on introspection tasks strengthens task-specific rationalization rather than privileged access.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.