acceptodds
Under review as a conference paper at ICLR 2027

Scenarist: Reshaping the Driving Data Distribution through Open-Ended Failure Scenarios, from Language to Policy

Abstract

Rare and safety-critical interactions are underrepresented in naturalistic driving logs and often cannot be collected deliberately, while the scenarios needed for improvement evolve with a driving policy's failures. We present Scenarist, a language-to-policy benchmark for configuring long-tail driving-data distributions through open-ended failure scenarios. Scenarist compiles natural-language failure-scenario intents into rule programs composed of reusable map, signal, and interaction primitives, with deterministic checks ensuring semantic and physical validity. It connects responsive traffic generation, safety-constrained response exploration, and trajectory-conditioned RGB synthesis to produce action-aligned supervision for vision-language-action (VLA) policies. Safe responses from Group Relative Policy Optimization provide positive training targets, while unsuccessful responses provide negative feedback. The benchmark contains 16,397 interactions across eight vehicle and pedestrian failure-scenario intents, including Scenarist-2K, a realistically rendered subset of 2,144 interactions. Our response policy reduces collision rate from 36.70% to 6.23% while substantially increasing safe-response diversity. VLA post-training with positive-sample SFT and negative-sample RL reduces collision rate from 24.96% to 1.42% and increases safe rate from 74.76% to 98.58% on Scenarist-2K validation, while maintaining general-driving performance on the NAVSIM benchmark. Thus, configurable synthetic supervision can improve long-tail safety while largely preserving routine-driving performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.