Plug, Play, or Re-Evaluate: Harness Components Across Models
Abstract
An agent harness is the system around a language model that provides the instructions, tools, context, and controls needed for the model to execute a task. Reusing components of this system for new models could reduce repeated testing and tuning, but comparisons of complete harnesses do not reveal which component matters or whether its effect carries over between models. We study each component by comparing the full harness with versions in which one component is removed, then examine whether a component's effect observed on existing models recurs on a new model. We discover three patterns. First, component effects are sparse across model–task–component combinations: removing a component leaves the recorded result unchanged in 84.3% of 1,120 combinations. Second, components exhibit model-specific compatibility: in a separate four-model analysis, the models whose results change differ across components in 68.3% of 360 model–task cases. Third, cross-model recurrence is limited: among 176 cases in which removing a component changes the target model's result, 94.3% have at most one other model showing a change. These findings show that an effect observed on one model should not be assumed to recur on another, motivating component- and model-specific validation before reuse. Source code is available at https://anonymous.4open.science/r/harness-transfer-artifacts-2041/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.