DocTrek: Tree-Guided Agentic Reinforcement Learning via Dense Rewards for Document Understanding
Abstract
Multimodal long-document question answering requires reasoning over hierarchical structures and heterogeneous content across many pages. Existing approaches discard such structures, use the hierarchy only as static retrieval units, or train the policy over flat sequences with trajectory-level feedback. We present DocTrek, which couples a section-structured document tree with token-position dense rewards for agentic reinforcement learning. A parsing–construction–integration pipeline converts raw PDFs into a tree whose section skeleton hosts multimodal elements as typed leaves, providing an addressable RL environment. Within it, we train the navigation-and-answering policy with a reward decomposition that converts page-level gold-evidence coverage into rewards placed at action-end tokens, providing dense credit for evidence retrieval and answer generation. Across four benchmarks, DocTrek improves the average score by 2.3 points over the strongest baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.