Procedural Faithfulness in Data Refinement: Can LLMs Execute Compositional, Order-Sensitive Recipes?
Abstract
Executing instructions that combine multiple operations requires following not only what to apply but also when. We study this procedural faithfulness in data refinement, where operation order and intermediate text states determine the correct output. Existing evaluations focus on individual tasks or end-to-end agent workflows, leaving the model's own ability to execute compositional, order-sensitive recipes insufficiently characterized. We introduce CDR-Bench, a diagnostic benchmark with 3,462 deterministic tasks across four domains and 29 operators, featuring paired execution-order and filter-placement variants. Across more than ten general-purpose LLMs, we find substantial gaps between individual-operation competence and complete recipe success in both rule-based and semantic refinement. Oracle interventions reveal difficulties in intermediate-text construction and filter-statistic computation, while code-assisted execution improves performance but leaves failures in recipe interpretation and order preservation. We develop training-free state-aware prompting that elicits explicit intermediate-state tracking and introduce Juicer, a specialized refinement model trained through broad operator learning followed by recipe-oriented adaptation. State-aware prompting improves compositional recipe success, while Juicer raises Order-M RS@3 from 7.02% to 31.82%. Overall, our study identifies procedural faithfulness as a distinct evaluation target in data refinement and provides a controlled framework for examining how models handle execution order, intermediate states, and stopping conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.