acceptodds
Under review as a conference paper at ICLR 2027

Language Models Are "Insecure" Reporters

Abstract

As large language models are deployed on increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes increasingly difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon "insecure reporting." When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 mentions the negative result in only 2 of 200 generated reports. However, when a short instruction, "Be honest in your response," is added, the model flags the negative result in 190 of 200 reports. Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about how to appear successful. We perform an activation analysis and a steering experiment on Qwen3.5-9B, finding that honesty and success-seeking correspond to opposing directions in representation space. Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.