Do Correct Answers Have Stable Latent Knowledge Support?
Abstract
Large language models are typically evaluated in terms of the correctness of their output. However, output correctness alone does not confirm whether the latent knowledge supporting an answer has stable internal representations. We therefore propose latent state stability (LSS) and knowledge topology stability (KTS) to investigate the stability of individual knowledge states and their relational structure, respectively. Specifically, we first separately average the model's internal representations of the same knowledge under each elicitation condition. Using the mean representation under the baseline condition as a reference, we remove shared shifts to obtain corrected representations. We then formulate LSS as the distance between each corrected representation and its reference representation, normalized against the typical distance between knowledge items of the same type. We formulate KTS as the cross-condition consistency of the ordered pairwise representation distances among knowledge items of the same type. Next, we construct EpiStableBench, which comprises 4,411 items spanning factual knowledge and mathematical reasoning, with 76,766 controlled queries. Across 16 LLMs from four families, representation stability varies substantially even among items answered correctly under the base condition. Within models, higher stability generally corresponds to higher accuracy under held-out elicitation conditions; across models, stability does not consistently increase with model size or overall accuracy. Meanwhile, we present LatentDPO to learn a knowledge-editing policy that accounts for the stability of latent knowledge. Experimental results show that integrating LatentDPO with multiple existing editing methods consistently improves editing success and generalization across different expressions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.