Harness2Policy: Harness Internalization for Long-Horizon Vision-Language Agents
Abstract
Long-horizon vision-language tasks require agents to preserve visual evidence, manage memory, and coordinate tool use. External harnesses guide these decisions, but successful-trajectory distillation can omit the control supervision and failure experience needed for autonomous execution. We present Harness2Policy, a framework that converts harness-guided exploration into process supervision. We search task-adaptive control configurations to guide tree-based exploration, retaining successful paths with explicit planning, memory, verification, and context-supported failure reflections. Using these trajectories, we build HarnessMind agents through supervised learning without harness-specific control prompts, enabling control through their responses and tool calls. Repeating this process with the updated model enables recursive co-improvement of the harness and the model. Across four benchmarks, HarnessMind (4B, 9B, and 35B) outperforms its base models with and without benchmark-specific oracle harnesses. On tasks whose reference runs exceed 20 turns, the 9B model improves accuracy from 7.4% with successful-path distillation to 14.3%. It also raises aggregate accuracy from 24.4% to 25.7% over its oracle-harness baseline with 31% fewer inference tokens. This work paves the way for future research on scalable self-improvement through the co-evolution of agentic harnesses, training data, and foundation models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.