DuoAgent: Self-Improving on Software Engineering from Visual and Structural Feedback
Abstract
LLM agents are increasingly used for real-world software engineering (SWE), yet they still struggle to resolve challenging practical software issues involving visual behavior, such as broken layouts or incorrect UI states, due to weak visual grounding, scarce training data, and inadequate supervision. To address these challenges, we propose a self-improving agentic framework based on visual and structure feedback to address multimodal SWE problems, where we ground agents in interactive browser environments, exploited during deployment, data generation, and training. In particular, we first propose DuoAgent, a with specialized agents: orchestrator, reproducer, solver, and verifier. The reproducer builds a browser sandbox and a fixed action sequence that triggers the reported defect. The solver analyzes the repository and proposes candidate fixes. For each candidate, the verifier rebuilds the software, replays the same actions, compares the resulting trace with the buggy one, and returns a visual and a structural diagnosis. This allows the solver to reason about the effect of its changes on the reported behavior. Then we propose a algorithm that localizes visual elements and injects perturbations into clean sandboxes with various strategies for data augmentation. We scale up the training data from 100 seed instances to 513 synthetic instances, each verified by re-executing the seed's action sequence on the mutated and fixed programs and checking that the injected defect is visible, stable, and removed by the inverse mutation. In addition, we design a novel post-training algorithm based on on-policy self-distillation, DuoOPSD, which trains a student model on its own rollouts with teachers conditioned on diagnoses of visual and structural discrepancies as privileged information. On SWE-bench Multimodal, DuoAgent establishes a new state of the art at 52.0%, substantially surpassing the previous best result of 36.0%. Meanwhile, training with DuoOPSD on our synthetic tasks significantly improves student models and consistently outperforms existing post-training algorithms trained on the same rollouts with the same number of optimizer steps. These results establish a new paradigm for self-improving multimodal SWE agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.