StratTR: One Strategy Does Not Fit All in Temporal Reasoning
Abstract
Temporal reasoning involves distinct operations, including event-graph reasoning for relation questions, date arithmetic for duration questions, and direct extrac- tion for factual questions. Existing systems, however, typically use one prompting scaffold or policy across these question types, making the reasoning procedure does not match the operation need for the question. We introduce StratTR, a trainable multistrategy agentic framework that support four strategies (graph con- struction, code generation, direct extraction, and hybrid reasoning). StratTR dis- tills routed, multi-call trajectories from a proprietary teacher into a single-call 4B policy. To train the strategy-conditioned policy, StratTR leverages the key deter- ministic temporal signals across heterogeneous artifacts with reinforcement learn- ing. Across seven benchmarks comprising 22,728 questions, StratTR with a 4B- scale backbone outperforms the strongest same-backbone baseline with macro- averaged exact match gains of 16.23%, top-performing API baseline 11.46%, and 3.04% for it’s teacher model with reduction of external model calls from 6.0 to 1.0. Further analysis suggests that learned verifiers and LLM judges can support candidate selection when deterministic checks are unavailable, although their ef- fectiveness varies across datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.