HITRA: HIERARCHICAL TRAJECTORY RERANKING FOR LONG-CONTEXT AGENT
Abstract
Reliable multi-hop question answering over thousands of documents requires agents to connect sparse evidence across long observation sequences under a limited active-context budget. Existing approaches largely follow a modular trajectory-processing paradigm, addressing memory, decision evaluation, and candidate selection. Although effective within their respective roles, they leave a coupled gap: compressed states can lose recoverable hypotheses, intermediate feedback can obscure distinct decision failures, and selection lacks a jointly trained ordering over partial trajectories. To address these issues, this paper proposes Hierarchical Trajectory Reranking for Long-Context Agent (HITRA), a framework designed to jointly address the interrelated challenges of memory compression, intermediate decision evaluation, and partial-trajectory selection. Especially, selective snapshot recall equips compact recurrent rollouts with recoverable prior hypotheses, allowing later evidence to correct state errors within a bounded active context. Hierarchical credit separately normalizes planning, action, and structural-validity feedback before combining local quality with terminal utility and exploration, distinguishing decision failures while aligning updates with task success. Listwise learning explicitly orders concurrent partial trajectories, with a separable surrogate that supports microbatch training and exactly matches the full listwise gradient at the behavior policy. Together, these methods enable evidence recovery, finer-grained attribution, and listwise ranking of concurrent partial trajectories. Extensive experiments demonstrate that across context sizes ranging from 50 to 6,400 documents, HITRA maintains (near-) optimal performance across all evaluated context scales. Controlled diagnostics further demonstrate improvements in memory recovery, root-cause localization, and trajectory selection: Compared to all baselines, the memory recovery rate increases by at least 8.7 percentage points (see Table 2 Recovery Rate), the root-cause localization rate increases by at least 9.7 percentage points (see Table 3 Root@1), and the optimal candidate selection capability improves by at least 18.0 percentage points (see Table 4 Hit@1).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.