acceptodds
Under review as a conference paper at ICLR 2027

MLDerive: Benchmarking Machine Learning Reasoning with Evidence-Guided Evaluation

Abstract

Solving machine learning (ML) problems requires large language models (LLMs) to apply domain knowledge through multi-step reasoning under problem-specific assumptions and constraints. Yet, how these models perform across diverse ML tasks and what specific reasoning errors drive their failures remain insufficiently characterized. We introduce MLDerive, a benchmark comprising 1,292 human-reviewed problems with reference solutions across various task types and ML knowledge areas. MLDerive links curated source problems to knowledge-aligned variants, enabling comparisons within source–variant families. To reliably assess model outputs, we propose an evidence-guided evaluator that combines tool verification, dependency-aware target checking, and semantic adjudication to determine whether solution claims jointly establish the requested conclusions, while retaining traceable evidence. Evaluation of six LLMs reveals task-level differences: all models achieve lower pass rates on optimization and algorithmic derivation than on proof tasks. Furthermore, qualitative analysis identifies three recurring reasoning failures: method misapplication, inferential insufficiency, and calculation-explanation discrepancies. Ultimately, MLDerive integrates quantitative performance, family coverage, and qualitative error analysis to provide a comprehensive framework for assessing ML reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.