Understanding Cross-Environment Agent Transfer through Task Semantics and Interaction Realization
Abstract
Teacher trajectories collected from executable benchmark variants provide a scalable source of supervision for agent post-training. However, given a target benchmark, it remains unclear how source benchmark variants should be selected. Source benchmarks may resemble the target in Task Semantics (TS)—including the task goal, domain, core capabilities, and success conditions—or in Interaction Realization (IR)—including how information is accessed, actions are executed, feedback is provided, and state is managed. We investigate which type of proximity is more predictive of effective transfer. To this end, we propose TIP, an analysis framework that characterizes benchmark relationships along the TS and IR dimensions separately. We generate and filter teacher trajectories from two groups of source benchmarks, fine-tune Qwen3 models via trajectory SFT, and evaluate transfer to a target benchmark. We find that sources more closely aligned with the target in interaction realization yield larger performance gains, whereas sources with greater task-semantic similarity consistently induce negative transfer across the evaluated models. Our study develops a benchmark-relation framework beyond task-semantic similarity, clarifies when trajectories generated by one benchmark can provide useful training signals for another, and offers testable guidance for source-data selection and the analysis of agent transfer across benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.