TaskMatch: Learning Task-Aware Relevance Beyond Semantic Similarity for Agent Harnesses
Abstract
Modern LLM agent harnesses augment the underlying model with reusable components such as external skills and procedural memories, making effective retrieval increasingly important as these repositories grow. However, existing retrieval pipelines often rank stored skills or memories by semantic similarity, which measures how similar a skill or memory appears to the current retrieval query but does not necessarily reflect how useful it is for completing the task. We introduce TaskMatch, a lightweight plug-and-play model that learns task-aware relevance from offline examples labeled as required, supportive, or irrelevant. Using the fixed query and skill and memory representations already available in the host retrieval pipeline, TaskMatch combines direct feature-wise matching with nonlinear interactions while leaving candidate generation and downstream agent execution unchanged. We instantiate the same architecture separately for two representative harness components—skill retrieval and procedural-memory retrieval—with fewer than 0.30M trainable parameters per model and no additional online LLM calls. Across both components, TaskMatch consistently improves retrieval ranking quality and downstream task execution. It raises SkillDAG success on ALFWorld from 67.10% to 81.43%, improves MemP Proceduralization success from 81.34% to 90.30%, and increases TravelPlanner hard-constraint satisfaction from 64.06 to 76.19. Our code is available at https://anonymous.4open.science/r/learned-relevance-scorer-agent-retrieval-4683/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.