Co-Harness: Collective Harness Evolution for Agentic Multimodal Large Language Model
Abstract
Existing multimodal large language model (MLLM) agents generally rely on predefined contexts, fixed tool-use policies, and static orchestration to support long-horizon reasoning and tool use, but these rigid execution mechanisms limit their ability to adapt to new and diverse multimodal agentic tasks. In this work, we aim to improve MLLM agents through harness evolution, which abstracts generalizable experience from successful and failed execution trajectories to refine model harness for related but previously unseen tasks. % To this end, we propose Co-Harness, a new framework that introduces collective learning into multimodal harness evolution to enhance the long-horizon reasoning and tool-use capabilities of multimodal agents. The core idea of Co-Harness is to jointly improve a set of specialized and complementary harnesses through shared experience and reciprocal feedback. Specifically, Co-Harness introduces two collective harness evolving methods: Collective Mixture-of-Harness Evolution (CoMixHE) and Memory-guided Collective Harness Co-Evolution (MemCoHE). CoMixHE collectively evolves specialized and complementary harnesses across and within task domains, expanding harness capacity and broadening the search space of harness optimization. MemCoHE jointly evolves the task model's harnesses and the optimization model's memory, enabling the optimizer to reuse past refinement experiences and adapt more effectively to changing failure patterns. % Extensive experiments on multiple multimodal agentic benchmarks demonstrate the effectiveness of Co-Harness and its transferability across tasks and models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.