More Than Correct Answers: Testing, Tracing, and Predicting Multi-Hop Execution in Transformers
Abstract
A correct answer to a multi-hop question does not establish that the model executed every required hop. The same answer may result from composing the supplied relations, recalling a stored association, or exploiting a prompt shortcut. We study multi-hop execution at three levels. First, we define execution by requiring predictions to remain correct under entity renaming and to follow counterfactual changes to every relation on the queried path. Tests on natural-language and Symbolic tasks reveal substantial failures hidden by high accuracy, including interference from familiar entity names and shortcuts that persist after renaming. Second, we trace execution within a single forward pass. State transplants show that middle-layer states carry an intermediate that later layers use to apply the next relation, while later states carry the completed answer. Component transplants identify recurring attention heads across tasks and benchmarks. Third, we predict execution from one unchanged prompt. Probes over layers identified by the transplant experiments outperform matched random layer ranges and generally outperform confidence- and correctness-based baselines, including under cross-benchmark transfer. Together, these methods provide a unified account of how to test multi-hop execution, trace it across depth, and predict it from a single forward pass.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.