HiToc: Hierarchical Token Compression Across the Vision-Language Pipeline for Long Document Question Answering
Abstract
Long-document question answering is computationally expensive due to long visual-token sequences. Compressing this context without disrupting document structure or losing cross-page evidence remains challenging. We propose HiToc, a training-free hierarchical token compression framework that progressively refines visual context across the vision-language pipeline. After visual encoding, Bottom-up Background Token Merging reduces background redundancy while preserving layout cues. During prefilling, Adaptive Page-level Token Assignment converts question-conditioned attention into soft page budgets and then selects tokens within each page, reducing attention-sink effects and preserving complementary cross-page evidence. During generation, layer-adaptive dynamic KV-cache compression reselects active visual context across decoder depths at each step. Across four document QA benchmarks, HiToc reduces the final active visual-token budget to under 20% of the original input while preserving 94.94–98.72% of baseline performance and achieving – LLM-stage speedups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.