Learn from Weaknesses: Student-aware Domain Specialization for Small Computer-Use Agents
Abstract
Small computer-use agents (CUAs) are attractive where cost, latency, or privacy matter, but they lag far behind large models in domain-specific software. We find that scaling up generic domain data yields limited gains, as small students differ in their weaknesses even within the same software. Existing failure-driven methods do adapt to the student, but they draw the reference for correct behavior from the student itself, leaving the failures it cannot yet solve uncorrected. We introduce LearnWeak, an annotation-free framework that makes specialization student-aware by contrasting the student with a stronger teacher, which can be a black-box API model. At the task level, it learns from correctable failures, tasks that the teacher solves but the student fails, and synthesizes practice that targets the missing skills. At the state level, it replays the student along the teacher's successful trajectories and corrects only the part of its action that deviates, depending on whether the failure stems from planning or execution. Training on only a few dozen trajectories per domain, LearnWeak improves EvoCUA-8B and OpenCUA-7B by 11.6 and 11.1 points on average across eight OSWorld domains, lets the 8B student surpass its 32B teacher on two domains where it initially trailed by more than 10 points, and outperforms existing data construction methods, including failure-driven training, under the same data budget. Our analyses further show that specialization data is student-dependent, as data built from another student's failures can even underperform the unadapted model. Together, these results suggest that specializing small CUAs hinges less on scaling data than on diagnosing and correcting the failures of the particular student.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.