acceptodds
Under review as a conference paper at ICLR 2027

How Well Do Models and Harnesses Work Together: Measuring Model–Harness Fit Beyond Task Scores

Abstract

LLM agents combine a model with an execution harness, so task scores reflect their joint behavior. These scores conflate model capability, harness-use capability, and task-solving support supplied by the harness. Because the same harness can impose different interaction demands on different models, we evaluate model–harness pairings through both task performance and operational validity. We introduce Valid Completion Rate (VCR) to measure normal termination with valid evaluator handoff, and model–harness Fit and Cross-Model Fit to summarize task performance achieved through valid completion. Using a common adapter that preserves native execution loops, we cross four general-purpose harnesses with multiple models across six benchmarks under matched conditions. Harness rankings vary across models and tasks, and task performance and valid completion can favor different pairings. Similar task scores can conceal distinct execution profiles, and low token use need not preserve outcomes. Cross-Model Fit identifies OpenCode as a stable aggregate leader within the evaluated panel, while model-specific preferences differ. These findings motivate joint assessment of task performance and valid completion alongside resource cost, training for transferable harness-use capability, and harness design adapted to the model and task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.