acceptodds
Under review as a conference paper at ICLR 2027

miniFlux: Exploring Agent-Directed Harness Evolution for Online Adaptation

Abstract

Agent harnesses have become a crucial factor in agent performance, shaping how models process context, take actions, and receive feedback. However, the support an agent needs can vary across tasks and evolve during execution. Existing approaches to automated harness evolution typically operate offline, relying on fixed task executions that may not capture diverse, emerging needs during deployment. We introduce miniFlux, an experimental framework for agent-directed online harness evolution within each task execution. Starting from a minimal harness (mini), the same model alternates between a main executor that advances the task and an evolver that adapts the active harness based on recent execution evidence. Triggered by periodic checks or executor requests, the evolver uses a bounded tool-use loop to revise runtime prompt guidance, tools, middleware, and context management. Component updates that pass interface and loading checks take effect through hot reload, preserving the progress without resetting the task. Evaluations on ARC-AGI-3, Tetris, and GDPval tasks show that miniFlux achieves higher mean task scores and improved action-level efficiency than the fixed mini baseline across evaluated models, and outperforms Codex in some evaluated settings. Cost analysis shows that online evolution adds relatively small overhead, representing less than 30% of the main execution cost in most evaluated settings. Evolution analyses illustrate that agents create and adapt task-specific components, such as analysis tools and validation middleware, to support information processing and task delivery. The ablation of evolution triggers shows that combining scheduled and agent-requested updates achieves higher mean performance than either mode alone. These findings suggest that adapting the harness online to emerging needs is a promising direction for improving agent performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.