acceptodds
Under review as a conference paper at ICLR 2027

-IF: A Multi-Variant Benchmark for Agentic Instruction Following in Executable Environments

Abstract

LLM agents must complete user goals while obeying standing and task-specific instructions. Their actions can alter state and activate different instructions. Existing agent benchmarks rarely pair executable training data with pre-specified instruction variants for testing this joint capability. To this end, we introduce -IF, a benchmark built on executable -bench environments. -IF releases 2,000 verified training programs with corresponding trajectories and provides two evaluation sets: Test evaluates 164 base tasks using one instruction variant each, whereas Test-Variants evaluates the same tasks under all five variants. Each instance combines a user goal, multi-turn interaction, executable tools, mutable environment state, and machine-scored instruction constraints. Across 23 eligible systems on Test, the best observed Task-and-Instruction Success is 40.24%, although Task and Instruction Success Rate scores reach 73.17% and 75.61%, exposing a joint-success shortfall. We then use -IF for an optimizer-paired study across three backbones. Full-trajectory SFT exceeds Fixed-Context GRPO and PPO by 5.08 and 4.80 points, while Online variants add and points over Fixed. -IF connects open training data, broad capability evaluation, and controlled experimentation in one reproducible platform, establishing verified trajectory supervision as a strong baseline while motivating better rewards and credit assignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.