acceptodds
Under review as a conference paper at ICLR 2027

Learning the Map, Not the Route: Curriculum, Capacity, and Transfer in Tool-Navigating Language Models

Abstract

We compare what language models learn from knowledge-graph navigation traces, which query a graph with tools and end in a machine-verifiable evidence path, and from free-form chain-of-thought distilled from a teacher (free-CoT). In a pre-registered study of 0.6B to 14B models on a synthetic, SHACL-governed kinship benchmark with strict format isolation, facts are withheld from the prompt and every edge in an emitted evidence path must have been retrieved during the episode, so a recited path scores zero. Each finding is tied to a registered criterion or labelled post hoc. (1) On the original corpus, navigation traces lose to free-CoT by 21.0 points at 12B. The decisive split inverts an ordering direction the trace compiler had held constant, and the navigation model's failures recite that constant: it learned the route, not the map. (2) When both corpora are regenerated with that direction varied, navigation leads by 23.7 points (Gemma-12B; seed-level interval +5.2 to +42.3), and the result replicates at Qwen3-14B under a rule registered before training (+13.9; +5.1 to +22.8). Varying it improves navigation by 25 to 29 points in every family tested; free-CoT drops in Gemma but not in Qwen3-14B. (3) Training on prose-only reasoning traces suppresses tool use (tool-call rates of 1.3% to 58.6% at 0.6B to 4B). Adding the tool calls back to the same prose restores tool use to 100%, and the gain from varying the direction holds for prose traces too (+23.8). At 4B the compiled notation adds 6.1 points on the original curriculum and 9.1 on the repaired one, a smaller effect that comes mostly from how the evidence path is written and does not carry over to transfer. (4) On a bijectively relabeled second domain, every model evaluated without in-context examples scores close to zero on the verified metric, for opposite reasons at the two sizes: the 4B model keeps the notation but loses the task, and the 12B model does the opposite. Two in-context examples fully restore the notation, and the navigation-trained 12B model then produces verified evidence on 65.1 5.9% of the new domain, 13.8 points above the untrained base model given the same examples (below the registered 15-point margin; the seed-level interval includes zero) and 5.8 points below the free-CoT model. The ability transfers, but no better than after free-CoT training. We report the null result with its repair, seed-level uncertainty next to every item-level interval, and three comparisons that would have favored our hypothesis but that our registered gates refused.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.