T-Spatial: Learning Utility-Aware Tool Use for Spatial Reasoning
Abstract
External tools are increasingly used to improve the performance of vision-language models (VLMs) on spatial reasoning tasks. However, executing a tool call successfully does not guarantee better performance. Inappropriate tool calls may even harm reasoning. Existing credit-assignment methods typically assign a single final-answer reward to an entire multi-turn trajectory, providing limited guidance on whether an individual call is beneficial, harmful, or redundant. This coarse-grained credit assignment makes it difficult to further improve the utility of tool use in long, multi-turn interactions. To address this challenge, we propose the Tool-Augmented Spatial Reasoning Agent (T-Spatial), trained with Difficulty-Aware Trajectory Synthesis (DATS) for supervised fine-tuning (SFT) and Tool Utility-based Policy Optimization (TUPO) for reinforcement learning (RL). DATS evaluates the target VLM on each training question and assigns a difficulty label based on its performance. According to these labels, DATS constructs a tool-assisted or a direct-answer trajectory for each question and teaches the target model to adaptively use spatial tools according to the difficulty of the question. TUPO combines sample-, trajectory-, branch-, and node-level tool-utility signals that are mainly computed from differences in answer scores across branches or before and after tool calls, allowing TUPO to reuse the existing answer evaluator without an auxiliary reward model. These complementary signals reduce the reward sparsity and make tool utility explicit at multiple granularity. Experiments on multiple spatial reasoning benchmarks show that T-Spatial improves spatial reasoning accuracy and tool-use utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.