acceptodds
Under review as a conference paper at ICLR 2027

DriveLab: Autonomous Research for Self-Driving Motion-Prediction Models

Abstract

Developing high-performing motion-prediction models for autonomous driving is an iterative and resource-intensive process: researchers must diagnose model failures, modify complex training and inference pipelines, debug implementations, launch expensive experiments, and prioritize next steps. We investigate whether this development loop can be automated using LLM-based autonomous research agents. To this end, we build an agent-operable environment around the Waymo Open Motion Dataset (WOMD) and study the autonomous improvement of two structurally distinct motion-prediction systems: Wayformer and MotionLM. We introduce DriveLab, an auto-research framework that organizes long-horizon exploration through three agentic roles atop a shared human-built infrastructure: Researcher Agents that propose high-quality, diverse hypotheses; sandboxed, skill-equipped Coders that implement changes and diagnose runtime failures; and an Evolution Manager that maintains persistent experimental memory and manages context. Under a fixed 24-hour compute budget on identical hardware, DriveLab achieves the highest Soft Mean Average Precision (Soft mAP) among the evaluated harnesses for all three base LLMs: averaged across base LLMs, it improves Wayformer by 9.5% and MotionLM by 13.0% over their starting checkpoints (up to 18.6% without a fixed time budget), and discovers non-trivial algorithmic improvements, such as new mode-aggregation mechanisms. We also identify the balance between research-idea novelty and code-execution stability as a key design consideration for auto-research and show that our method handles both well. Our results provide strong evidence that autonomous LLM agents can take on substantial parts of the development of complex driving-behavior models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.