acceptodds
Under review as a conference paper at ICLR 2027

From Failures to Harness: Admission-Gated Runtime Patch Selection for Hard Harness Defects in LLM Agents

Abstract

Large language model (LLM) agents rely on a runtime harness for tool calls, execution budgets, and recovery. Evaluating repairs to this layer requires distinguishing task success from recovery under an exercised defect. We introduce an injected-defect setting and an evaluation protocol that measures repair under recorded trigger exposure, same-family transfer, and clean-task regression. TARPS (Trigger-Aware Replay for Patch Selection) implements candidate search and admission using replay and reference-repair screening. Across seven frozen domain–defect cohorts on τ²-bench, reward-only selection raises mean repair from 10% for immediate first-candidate selection to 57%; adding clean-task checks reaches 76%, matching full TARPS. TARPS attains 76% transfer versus 71% for reward plus clean checks; the incremental effect of its remaining checks is unresolved at this sample size. Separate admission analyses show that reference-mechanism inclusion, refusal, and transfer measure distinct properties, and coverage changes which failures enter evaluation. Together, these findings identify execution checking as the main source of the observed gain and clarify what admission and trigger-aware evaluation measure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.