acceptodds
Under review as a conference paper at ICLR 2027

Bandit Limited Discrepancy Search for Agent Harness Configuration

Abstract

Coding-agent harnesses decompose work into a task DAG and dispatch each node to a sub-agent, but they fix the per-node configuration—agent type, model, reasoning effort, turn budget—from a plan the language model emits once, and never revisit it. We formalise the resulting problem as agent-DAG configuration selection and introduce BLDS-AF, a discrepancy-limited search over the Hamming ball around that plan that replaces the usual independent-arm bandit with a factored surrogate: a ridge posterior over node-arm and edge-pair indicators, weighted by compute allocation, with a group-wise prior that lets the model collapse onto an additive basis while data is scarce and switch interactions on as evidence accumulates. Every evaluation is pooled across configurations rather than treated as evidence about one configuration alone, and a composed candidate from the surrogate's own argmax is evaluated like any other. BLDS-AF beats the deployed greedy cascade and all fourteen baselines we tried, including Hyperband, TPE, evolutionary search, and UCT; it also beats plan-local coordinate descent by a significant margin, an advantage that grows monotonically with the size of the configuration space and holds at every interaction strength. It beats every plan-blind method on NAS-Bench-Macro with every interval excluding zero, and reaches zero normalised regret at a modest budget on a DAG of live multi-turn coding agents. The gains come from pooling statistical strength across configurations: a pre-flight diagnostic derived from the bias–variance condition of the factored surrogate predicts in advance on which substrates this works and on which it does not.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.