Discover-Many-Worlds: Reinforcement Learning for Scientific Discovery in Automatically Generated Physics
Abstract
Scientific discovery in physics requires a process of designing experiments, revising hypotheses, and inferring unfamiliar laws from observations. We introduce Discover-Many-Worlds, a reinforcement learning framework for training scientific discovery agents in automatically generated physics worlds. Our world-generation pipeline combines LLM-driven construction, procedural simulation, and verification testing to generate and validate diverse environments at various difficulty levels with no human intervention. We train a 4B-parameter Qwen model with GRPO, increasing its capabilities on DiscoverPhysics up to the performance level of Claude Sonnet 4.6. We also observe improved performance on other scientific and reasoning benchmarks. We demonstrate that the RL-trained agent engages in more extensive experimentation, where it gathers more observations before committing to an explanation and considers a broader range of physical laws instead of defaulting to familiar interactions. These results demonstrate that automatically generated physical worlds can provide effective training environments for experimentation and reasoning about physical phenomena, both of which are essential for scientific discovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.