LUDKA Bench: Can AI Agents Solve Nonstandard Machine Learning Problems?
Abstract
Solving challenging, nonstandard problems in artificial intelligence and machine learning requires identifying useful structure, choosing an effective approach, and translating it into a working solution under limited time and experimentation. We introduce a benchmark comprising 20 of the most challenging tasks from international AI olympiads and national qualifying competitions from multiple countries to evaluate these capabilities in language models and agentic systems. Beyond measuring independent problem solving, the benchmark asks whether providing the right direction is sufficient for a system to solve a difficult task. We evaluate three levels of assistance: no guidance, a strategic hint, and a detailed description of an approach used in a solution ranked among the top three on the original private leaderboard. In our current evaluation, the strongest tested system achieves 18.5 mean rank without guidance and 8.0 with a detailed description of a reference approach. Competition-experienced human participants achieve a mean rank of 6.09 under the same evaluation conditions with detailed solution descriptions, on tasks they had not previously solved. Additional guidance improves aggregate performance, but even full descriptions leave a gap relative to these participants. These findings show that even detailed guidance based on a successful solution is often insufficient for agents to achieve a comparable result without further human assistance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.