acceptodds
Under review as a conference paper at ICLR 2027

ProcIF: Evaluating Procedural Instruction Following in Large Language Models

Abstract

Large language models increasingly execute tasks that specify not only what outcome to produce, but also how evidence should be collected, actions should be ordered, failures should be handled, and results should be validated. A correct final answer, however, does not establish that the prescribed process was followed; conversely, procedural compliance does not guarantee a correct outcome. We introduce ProcIF, a cross-domain, two-track procedural instruction following benchmark. It contains 342 instances and 4,157 executable process constraints. Track A adds task-specific procedural constraints to writing, scientific reasoning, workflow synthesis, and tool-use tasks, whereas Track B converts source harness trajectories into visible, observation-conditioned Skills or standard operating procedures. We report strict Procedure Adherence Score (PAS), source-aware Task Success Rate (TSR), and their instance-level conjunction. Across the evaluated models, the best strict PAS and Joint scores reach only 41.23% and 35.96%, respectively. These results show that models often satisfy local requirements while failing the complete procedural contract, making procedural instruction following a challenging and substantially unresolved capability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.