acceptodds
Under review as a conference paper at ICLR 2027

FACT: Feedback Attribution with Controlled Trajectories for Long-Horizon Artifact Agents

Abstract

Long-horizon artifact agents transform persistent environments through many interdependent tool actions before receiving feedback on the completed artifact. This delayed and often subjective feedback makes it difficult to determine which earlier decisions improved the outcome and to turn that evidence into reliable policy updates. PPO and group-relative policy optimization methods such as GRPO assign rollout-level or group-relative advantages to sampled sequences; delayed terminal feedback and stochastic tool execution leave action-level credit unresolved. We therefore introduce FACT, a critic-based PPO formulation that defines action credit by the expected effect of a committed action on terminal preference. FACT executes alternative actions from the same restored workspace and continues each branch under a shared policy, yielding matched long-horizon comparisons. These comparisons supervise an action-conditioned transition critic for ordinary rollouts, while held-out interventions calibrate the resulting updates. We further introduce LongEditBench, a source-grounded benchmark for long-horizon video editing and contextual revision under follow-up user feedback. Experiments show that FACT improves completion-adjusted artifact quality and contextual revision, while controlled evaluations demonstrate more faithful action attribution. Consistent gains across transfer benchmarks further validate its effectiveness across persistent artifact environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.