Learning Precise Spatiotemporal Relational Reasoning in Large Language Models
Abstract
Inventory tracing and historical audits require complete answers about which objects reached a destination, when they shared a location, and what incomplete observations establish. These queries combine changing spatial dependencies, temporal aggregation, and evidence limits; a plausible item or a correct final location alone may not complete the request. We study how Large Language Models can learn these joint reasoning requirements. Executable task construction supplies questions, answers, and derivations for both evaluation and supervision. To scale language supervision, we use a strong teacher's interactions on 9,000 seed instances to train two smaller question and reasoning polishers. Applied to newly generated histories, they yield 85,940 accepted training records under source and rendering constraints. We examine learning with smaller corpora and follow the expanded-corpus solver through supervised fine-tuning and reinforcement learning. Recorded Qwen3.5-0.8B validation accuracy starts at 0%, reaches 38% at the supervised-to-reinforcement-learning handoff, and reaches 45% for the selected RL model. Paired answers show corrected temporal and occupancy results alongside persistent evidence-state and value errors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.