acceptodds
Under review as a conference paper at ICLR 2027

From One Query to Many Tools: Latent Demand Decomposition for Multi-Tool Retrieval

Abstract

Modern agentic systems depend on selecting relevant tools before a language model can execute a task. As tool registries grow, selecting a small candidate set becomes a retrieval problem. However, single-embedding retrievers can struggle when a query contains multiple semantically distinct demands, as a single query representation may not align closely with every required tool. We introduce LatDec (Latent Demand Decomposer), a retrieval architecture that maps a query to multiple latent demand representations in a single encoder forward pass. Its independent heads are trained using Hungarian matching to align each representation with a distinct target tool. At inference, Confidence-Proportional Head Budgeting (CPHB) converts each head’s confidence into its share of the retrieval budget. As a complementary signal, LatDec constructs a confidence-weighted population vector from all head representations, blends it with the frozen base query embedding, and combines the resulting ranking with the CPHB ranking using reciprocal rank fusion. We evaluate LatDec on two benchmarks: CapCorpus, a 22K-tool MCP retrieval benchmark that we construct and validate through independent human judgment, and the independently developed LiveMCPBench benchmark. On CapCorpus, LatDec substantially outperforms the frozen Sentence Transformer, matched bi-encoder, and multi-vector baselines. On LiveMCPBench, the benchmark’s current sample size does not support a clear distinction between the full LatDec model and the baselines. Within LatDec, however, adding embedding-space fusion increases retrieval performance over CPHB alone by 26.3% relative, corresponding to an absolute Recall@10 improvement of 0.0797 (95% CI [0.0445, 0.1172]). In a quartile analysis, the gain is larger for queries whose confidence-weighted population vector is least aligned with the base query embedding, highlighting the importance of balancing demand-specific specialization at the head level with preservation of the query’s overall semantics at the population level.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.