How Well Do Models and Harnesses Work Together: Measuring Model–Harness Fit Beyond Task Scores
Abstract
LLM agents combine a model with an execution harness, so task scores reflect their joint behavior. These scores conflate model capability, harness-use capability, and task-solving support supplied by the harness. Because the same harness can impose different interaction demands on different models, we evaluate model–harness pairings through both task performance and operational validity. We introduce Valid Completion Rate (VCR) to measure normal termination with valid evaluator handoff, and model–harness Fit and Cross-Model Fit to summarize task performance achieved through valid completion. Using a common adapter that preserves native execution loops, we cross four general-purpose harnesses with multiple models across six benchmarks under matched conditions. Harness rankings vary across models and tasks, and task performance and valid completion can favor different pairings. Similar task scores can conceal distinct execution profiles, and low token use need not preserve outcomes. Cross-Model Fit identifies OpenCode as a stable aggregate leader within the evaluated panel, while model-specific preferences differ. These findings motivate joint assessment of task performance and valid completion alongside resource cost, training for transferable harness-use capability, and harness design adapted to the model and task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.