acceptodds
Under review as a conference paper at ICLR 2027

TIDE: Targeted Probe-Guided Interventions for Cross-Model Knowledge Recovery

Abstract

Large language models (LLMs) are increasingly deployed for complex, multi-step tasks, making the reliability of their generated answers ever more important. Yet, even when an LLM internally computes the correct answer, it may fail to express that knowledge in its output, resulting in a knowledge-prediction gap in which the model internally “knows” more than it says. Bridging this gap is critical for improving the reliability of LLMs, particularly in high-stakes scenarios where incorrect or unfaithful responses can have severe consequences. However, existing approaches face two fundamental challenges. (1) Most are model-specific: they learn to extract answers from the representations of a single model, potentially relying on model-specific features rather than capturing the task-level information that generalizes across models. (2) Recovering latent knowledge with a probe does not directly improve the model's generation; an effective mechanism is needed to transfer probe predictions back into the model's output with minimal performance loss. To address these challenges, we propose TIDE (Targeted Probe-Guided Interventions for Cross-Model Knowledge Recovery), a framework that bridges model-internal knowledge and generated answers through cross-model probing and targeted intervention. TIDE trains a single probe across multiple models for a given task, enabling it to capture task-level information shared across models rather than model-specific representations and thereby generalize to unseen models. TIDE then identifies key attention heads involved in writing answers and selectively injects the probe's predictions into these heads, allowing the inferred answer to flow through the model's native output-writing mechanism. Extensive experiments across multiple LLMs demonstrate that TIDE consistently outperforms existing baselines, showing that a single cross-model probe can capture transferable information and effectively narrow the gap between internal and generated answers. Our code can be found at https://anonymous.4open.science/r/TIDE-D338/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.