How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models
Abstract
A model's reasoning trace can expose errors and misbehavior absent from its final response, yet closed-source deployments may show users only that response and a short reasoning summary. Here, we ask how much monitoring performance this reduced access preserves and introduce an observability ladder that holds completed runs fixed and varies access to prompts, responses, reasoning summaries, full traces, and internal model features. We use final-answer correctness as an externally labeled test case across five open-weight models and three benchmarks, with each model writing a self-summary from its completed trace alone. Linear classifiers trained on text features and embeddings reach mean AUROC 0.73 from the prompt alone and 0.75 with the response. The summary then adds 0.02 (95% question-bootstrap CI [0.01, 0.03]), the full trace a further 0.04 [0.03, 0.05], and internal features 0.03 [0.02, 0.04], reaching 0.84. However, a summary-level reader given the trace's word and sentence counts without its text trails the same reader given the full text by 0.01 [0.01, 0.02], so part of the trace's gain is length. Without the prompt, the summary's gain jumps to 0.16 [0.13, 0.18], consistent with the summary standing in for the missing question, and for the linear classifiers, telling a correct from an incorrect run of the same MMLU-Pro question barely rises above chance (0.53 to 0.57). Much of the across-question performance, then, is attainable from the prompt alone, without observing the particular run. Two LLM readers recover far more within-question discrimination from summaries and traces than the linear classifiers when the prompt is withheld, yet reach their highest concordance from the prompt and response alone when it is visible, so what a summary reveals depends on the reader. We have measured one property and one kind of summary, but because the ladder holds the run fixed and varies only what the reader sees, it can be applied to misbehavior with independent labels, to provider summaries, and to stronger readers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.