acceptodds
Under review as a conference paper at ICLR 2027

Implementation-Specific Effects in Language-Model Harnesses

Abstract

Runtime code determines which context a language model receives, how often it is called, and when its answers are revised. HarnessBench studies these choices through factorial ablations of context management, memory, feedback, and planning. Each component has a shared null and two concrete active implementations. A complete-output replication contains 4,608 evaluations across three fixed model identifiers, two task domains, 16 configurations, and both implementation variants. Direct paired contrasts show differences in marginal effects between global implementation bundles, including Planning on math and code and Context and Feedback on code. These contrasts jointly change every active component and do not isolate a single component implementation. The direction and statistical conclusion also vary across models, domains, and endpoints. Near-call-budget controls, repeated runs, and a 500-task MATH-500 validation show that these treatments are implementation bundles whose effects include inference cost and operational reliability. On two local models, 3,200 planning episodes show no clear accuracy recovery after summary truncation is removed, despite substantially higher input-token costs. An isolated 1,600-condition feedback experiment shares first answers across revision prompts: self-checking outperforms a same-call neutral revision on Llama but not Qwen. HarnessBench therefore supports implementation-level diagnosis; two implementations per component do not estimate population-level implementation variance or identify a universal best harness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.