STQA: A Dataset and Road-Subgraph-Augmented Language Model for Spatiotemporal Traffic Question Answering
Abstract
Spatiotemporal question-answering is crucial in transportation scenarios. However, existing traffic models typically focus on fixed numerical tasks, while general-purpose language models struggle to integrate temporal dynamics, directed road network dependencies, and heterogeneous contextual evidence. We investigate whether diverse traffic-analysis tasks can be unified through a natural-language interface while retaining structured spatiotemporal evidence. To this end, we introduce STQA, a dataset of approximately 120,000 question–answer pairs spanning ten subtasks across forecasting, imputation, classification, and explanation. STQA aligns traffic observations with road-network context to support both numerical analysis and contextual explanations of traffic patterns and event responses. We further propose RoadSubLM, a road-subgraph-augmented language model that uses a directed graph encoder to represent local road-network context as compact graph soft tokens. Auxiliary supervision reconstructs adjacency and road distances from graph-token hidden states, encouraging road-network structure to remain recoverable after graph–language integration. On STQA, RoadSubLM outperforms the evaluated language-model and time-series-language baselines across all ten subtasks, with particularly large gains in imputation, spatial-propagation classification, and contextual explanation. The dataset and code are available at https://anonymous.4open.science/r/STQA-66F3.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.