acceptodds
Under review as a conference paper at ICLR 2027

PersonalArc: Harnesses at Both Ends of Agent Learning

Abstract

Personal-context agents must answer from task-scoped evidence while learning from prior executions. This creates two linked challenges: controlling how evidence enters a training trace, and deciding whether a learned answer should replace the base answer. We introduce PersonalArc, a framework that places restricted harnesses at both ends of agent learning. A task-conditioned harness declares stage inputs, artifact types, tool scopes, and resource limits before execution. Workspace-grounded mutations are audited and distilled into content-only stage supervision. At inference time, a back-end fusion harness governs evidence access, rubric-blind composition, and final-answer selection. It adopts the fused candidate when both comparison orders prefer it and numeric retention passes, returning the base answer exactly on rejection. Together, the two harnesses link experience acquisition and answer deployment, turning audited executions into reusable supervision while keeping final-answer adoption under explicit control. We evaluate the full evaluation sets of four benchmarks, with three generating backbones in the main comparison. Averaged across these backbones, PersonalArc improves Yes Rate and Coverage by 27.0% and 18.2% relative to Base Search, respectively, while increasing F1 by 6.77 points. It also outperforms the evaluated aggregation baselines at every tested token ceiling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.