MERIDIAN: A Long Context, Robust Instruction Following Benchmark
Abstract
We introduce Meridian, a benchmark for long-context (LC) instruction following (IF) over real-world documents. IF is the model capability that allows users to condition model output. Our goals are to enable diagnostics of model failures for this setting by exploring robustness to instruction positions in the prompt and resistance to instruction injection. We design prompts that state document-understanding tasks together with a set of typed constraints incremental to the task. We target one general knowledge domain and four domains relevant to knowledge work (finance, legal, medical, cloud service documentation) both through documents and constraints specific to the domains. We validate instruction prompts via human review. We report Task score, and IF score for the instruction prompts, and a weighted combined score. We test IF across changes in context length, instruction count, and instruction placement; we further explore single- and multi-turn settings. We evaluate current models and show that 1) Combined scores on average degrade (from to ) when context length increases from 8k to 128k tokens, and are sensitive to prompt position; 2) IF score has headroom (max of for the most complex prompts), and is only moderately correlated with writing quality. We plan to release the benchmark and code upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.