Beyond Answer Correctness: Executable Derivation Faithfulness for Urban Spatial Reasoning
Abstract
Urban spatial reasoning typically relies on external tools to acquire spatial evidence and on multi-step derivations to reach final decisions. However, answer correctness does not imply derivation faithfulness: models may produce correct answers despite misusing spatial evidence or making errors in intermediate derivations. Outcome-only evaluation and training cannot distinguish these trajectories from faithful ones that yield the same answer. To address this issue, we propose UrbanSDF (Urban Spatial Derivation Faithfulness), a framework for derivation faithfulness in urban spatial reasoning. UrbanSDF explicitly models how entity bindings determine tool arguments, how returned observations support spatial computations, and how these computations determine final decisions. To verify these dependencies, we construct executable derivation certificates that assess derivation faithfulness separately from outcome correctness. The resulting feedback guides a progressive training strategy: after initial tool-protocol supervision, we address detected errors through targeted process supervision with faithful demonstrations, improving the model's ability to generate faithful trajectories under limited sampling. Building on this supervision, we employ Conditional-Dual GRPO optimization during online interaction, using separate outcome and process feedback to favor faithful derivations among answer-correct trajectories. Experiments on four types of spatial reasoning tasks show that UrbanSDF maintains task accuracy while substantially increasing the proportion of faithful trajectories among those yielding correct answers, with an average improvement of 45.6% over the strongest baseline. Our code and data are available at https://anonymous.4open.science/r/UrbanSDF-5A02.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.