acceptodds
Under review as a conference paper at ICLR 2027

Workflow-Style Instruction Following with Atomic Synthesis and Adaptive RL

Abstract

Complex instruction following is almost always scaled along one axis—more constraints on a single task—yet many instructions are workflows: interdependent sub-tasks under conditional branches, the hardest of which turn on the model's own intermediate output, so the correct route cannot be known before the response is written. Such execution flow is largely absent from existing data, and checklist-style verification does not match it: a checklist is fixed before the response, whereas the path to be checked is not. We propose CASCADE, an iterative synthesis framework that expands a seed instruction into a workflow-style complex instruction by adding exactly one atomic unit per round—a constraint, a sub-task, or a conditional branch. Verification follows the same structure: a check graph models an instruction's execution as a directed acyclic graph, and each response is judged node by node along the path it ought to follow. The same pipeline on held-out seeds yields CascadeEval, a benchmark of such instructions. Because expansion is atomic, synthesis also yields, at no extra cost, a difficulty spectrum for every instruction: the ordered sequence of versions from the seed to the hardest one. We exploit it with SPIRAL, a reinforcement learning (RL) data-selection algorithm that chooses, at each step, which version of each instruction to rollout, keeping training matched to the model's current ability. On Qwen3.5-9B, RL on the hardest CASCADE instructions raises the instruction-following average by +11.00 points over the strongest existing corpus; SPIRAL adds a further +1.88 points over that setting while cutting wall-clock training time by 40%. Our data and code will be made publicly available after peer review.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.