SOLVE: Adaptive Evidence Acquisition for Verifiable Multi-Agent Decision Making
Abstract
Large language model (LLM) agents increasingly integrate tools and multi-step reasoning, but consequential decisions may often require selective evidence acquisition, non-compensatory constraints, and auditable termination. We formulate this setting as evidence-constrained decision making and introduce MarsSite-Bench, a 600-task benchmark spanning six task types with explicit evidence budgets, structured outputs, and method-independent constraint-gated evaluation. We further propose Sequential Orchestration with Logical Verification and Evidence Acquisition (SOLVE), which coordinates adaptive evidence acquisition and role-specialized reasoning through a shared decision state. On 120 hidden-test tasks, SOLVE achieves a Primary CGDQ of 0.6488, compared to 0.1795 for ReAct. Matching ReAct’s evidence access raises its score to 0.3976 but leaves a substantial gap. A real-Mars case study further demonstrates capabilities in regional ranking and footprint-level abstention under heterogeneous geospatial evidence. These results support adaptive acquisition and constraint-aware orchestration for verifiable LLM-agent decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.