Don't Wear Too Many Hats: Capability Modularization for Multi-Document Analytical Question Answering
Abstract
Multi-document analytical question answering (MDocAQA) requires coordination among different information processing capabilities over evidence distributed across files. In existing methods, individual agent typically needs to perform operations across multiple capabilities. Since single agent often struggles to sustain effective reasoning and reflection throughout long-horizon workflows that require coordinating diverse capabilities, its performance tends to be poor. To address this issue, we propose an MDocAQA framework called Sensemaker. Inspired by the model of expert sensemaking, which represents information analysis as an iterative process containing distinct information processing activities, Sensemaker organizes MDocAQA into three layers (i.e., task layer, capability layer, and tool layer) and explicitly defines five modules responsible for different information processing capabilities in the capability layer. During task execution, the task-layer agent dynamically coordinates these modules, each instantiated as a specialized agent equipped with dedicated tools. This explicit capability modularization reduces the workload on individual agent and makes its intermediate states easier to inspect. Extensive experimental results on various benchmarks show that Sensemaker outperforms representative MDocAQA methods, with up to a 29.71% relative improvement in average answer score over the strongest baseline method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.