Lego-UX: Unlocking the Value of Human-Agent Interaction Trajectories
Abstract
Recent advances in Large Language Models (LLMs) have substantially improved coding agents on single-turn benchmarks such as SWE-bench. However, real-world coding assistance is naturally interactive: users refine requirements, report failures, and correct unsuccessful attempts across turns. In this setting, user experience (UX) encompasses both test-verified task completion and interaction quality, including how effectively the agent follows and adapts to user feedback. Improving this experience calls for learning from real-world human–agent trajectories, which provide valuable supervision for multi-turn interactions but are difficult to leverage reliably. Their quality varies widely, complicating training data selection, while their underlying tasks often lack reproducible environments for verification and replay. We introduce Lego-UX, a unified framework that uses interaction-aware trajectory selection and evidence-bounded task reconstruction for training and multi-turn evaluation. It identifies reliable supervision using behavioral and execution signals, while recovering executable tasks that facilitate multi-turn evaluation and controlled training rollouts. Using the reconstruction pipeline, we construct Lego-UXBench, a 95-task benchmark for multi-turn agent user-experience. In the experiments, we evaluate 11 frontier models on Lego-UXBench and observe that similar task-completion performance can correspond to markedly different interaction quality. Fine-tuning Qwen3.5-35B-A3B on selected trajectories improves both measured dimensions and raises the unweighted eight-metric macro score across four benchmarks by 15.2 points (7.75 to 22.91).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.