acceptodds
Under review as a conference paper at ICLR 2027

Recovery Under Drift: When Failed Context Helps or Hurts

Abstract

Failed context can contain both misleading content and correct task information, so removing it changes two things at once. We study this trade-off with controlled interventions that vary correct facts inside a partly wrong record, correct facts supplied separately, failed-code visibility, the source of failures used for training, and conversation retention at fixed files. In tool-use and ALFWorld recovery states, adding correct facts to the retained record raises its benefit for Qwen3-8B and Gemma-4-12B-it. A complete current reference lowers retention benefit in the tool setting, whereas the ALFWorld intervals do not establish the same change. A secondary Qwen tool comparison finds that partial external facts can first increase the record's benefit before a complete reference lowers it. Code experiments show that measured recovery also depends on restart context and failure source: for one 7B model, omitting failed code and feedback raises success within 64 attempts from 22.1% to 55.5%, and matched-source training gains exceed mismatched gains across five model settings. On 34 MettleBench failures at identical repository files, retaining the pre-checkpoint conversation raises mean recovery from 11.8% to 35.8%. Recovery evaluations should state what each condition retains, compare repair with fresh starts, and match inference budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.