MLDerive: Benchmarking Machine Learning Reasoning with Evidence-Guided Evaluation
Abstract
Solving machine learning (ML) problems requires large language models (LLMs) to apply domain knowledge through multi-step reasoning under problem-specific assumptions and constraints. Yet, how these models perform across diverse ML tasks and what specific reasoning errors drive their failures remain insufficiently characterized. We introduce MLDerive, a benchmark comprising 1,292 human-reviewed problems with reference solutions across various task types and ML knowledge areas. MLDerive links curated source problems to knowledge-aligned variants, enabling comparisons within source–variant families. To reliably assess model outputs, we propose an evidence-guided evaluator that combines tool verification, dependency-aware target checking, and semantic adjudication to determine whether solution claims jointly establish the requested conclusions, while retaining traceable evidence. Evaluation of six LLMs reveals task-level differences: all models achieve lower pass rates on optimization and algorithmic derivation than on proof tasks. Furthermore, qualitative analysis identifies three recurring reasoning failures: method misapplication, inferential insufficiency, and calculation-explanation discrepancies. Ultimately, MLDerive integrates quantitative performance, family coverage, and qualitative error analysis to provide a comprehensive framework for assessing ML reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.