acceptodds
Under review as a conference paper at ICLR 2027

Bifrost: Steering Strategic Trajectories to Bridge Contextual Gaps for Self-Improving Agents

Abstract

Autonomous agents excel in self-improvement through reflection and iterative refinement, which reuse successful task trajectories as in-context examples to assist subsequent reasoning. However, shifting across tasks often introduces a context mismatch. Hence, existing approaches either discard the trajectories or manipulate them using heuristics, leading to a non-negligible fine-tuning cost or unguaranteed performance. To bridge this gap, we reveal a context-trajectory correlation, where shifts in context are highly parallel with shifts in trajectory. Based on this finding, we propose BrIdge contextual gap FoR imprOvised trajectory STeering (Bifrost), a training-free method that leverages context differences to precisely guide the adaptation of previously solved trajectories towards the target task, mitigating the misalignment caused by context shifts. Our trajectory adaptation is conducted at the representation level using agent hidden states, ensuring trajectory transformation accurately aligns with the target context in a shared space. Empirical experiments reveal that Bifrost consistently surpasses existing trajectory reuse and fine-tuned self-improvement methods by 2%–14% across adaptation tasks in math problem solving (AQUA to GSM8K), question answering (ARC to GPQA), and code generation (HumanEval to LiveCodeBench), demonstrating that agents can effectively leverage past experiences despite substantial context shifts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.