acceptodds
Under review as a conference paper at ICLR 2027

Anchoring the Bridge Across Hops: Unlocking Latent Multi-Hop Reasoning under Knowledge Injection

Abstract

Test-time training (TTT) keeps parametric knowledge current by writing updated facts into the parameters, but those facts are queried in multi-hop form at inference and their combinations cannot be enumerated during training. Whether a model trained only on the atomic form composes arbitrary chains within a single forward pass remains a critical question. We evaluate this on a two-hop benchmark of updated facts in which no training sample contains both hops of a chain. We show that latent composition remains hard for existing rewriting methods for knowledge injection (paraphrase, keyword diversification, and added context), which help the model acquire the updated facts but leave composed accuracy low (below at B). We attribute this to the two hops never being aligned on the bridge entity during training. We propose Bridge Anchoring by Common Enrichment (BRACE), which reuses the same descriptor pool on the bridge in both hops, so that the two hops present it in a common form. Across four pretrained LLMs spanning two families and B–B and five domains, BRACE improves composed accuracy over the strongest baseline in all model and domain cells, by to in relative terms, and raises the composition rate, composed direct-QA accuracy restricted to the chains that answer at least half the paraphrases of both hops, up to . A count-matched control isolates that alignment: at matched atomic acquisition, drawing the bridge descriptor from the same pool in both hops raises composed accuracy, and a type-incorrect descriptor costs acquisition rather than composition. A logit lens reads BRACE resolving the bridge to layers earlier. We further assess two other routes to elicit composition, chain-of-thought prompting and composition demonstrations on held-out entities, and find that the first lowers composed accuracy by recalling the pretrained bridge, while the second raises composition only when it precedes the target atoms and only on BRACE's training text. This paper presents a basis for compositional generalization under continual knowledge injection, where the chains queried at inference are by construction absent from the training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.