ARCUS-X: Adaptive Rule Compliance Under Sequential eXecution for Long-Horizon Language Model Evaluation
Abstract
Long-horizon language-model evaluations often reduce performance to final-answer accuracy, obscuring failures that accumulate across dependent sequential transitions. We introduce ARCUS-X, an auditable benchmark and evaluation protocol for decomposing deterministic sequential execution into trajectory correctness, continuity, rule binding, horizon control, response validity, and excess generation. ARCUS-X uses procedurally generated toroidal-grid tasks with exact per-transition ground truth, controlled lexical and transition perturbations, raw-stream preservation, and deterministic post-run diagnostics. Across 13 models, ARCUS-X exposes substantial separation between aggregate trajectory quality and observed failure signatures. In the adaptively sampled runs, GPT-5-mini attains 84.8% step accuracy, 42.5% exact match, and 84.5% continuity, but 25.9% generation efficiency and 11.2% invalid responses. The valid-error partitions differ sharply: state-tracking signatures account for 93.8% of Qwen 3 Coder 480B A35B probes and 84.3% of GLM-5 probes, transition-rule signatures are largest for Mistral Large 2512 (23.6%), and semantic-interpretation signatures are largest for Kimi K2.5 (23.2%). Sampled fracture behavior also varies by perturbation tier. Because AFS allocates horizons from observed performance, pooled aggregates describe the acquired support rather than a common uniform horizon distribution. ARCUS-X therefore provides a reproducible diagnostic record of explicit state-transition execution rather than a universal model ranking or measure of general long-horizon agency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.