SIRRA: Matched Evaluation of Standing-Instruction Reapplication after Temporary Non-Applicability
Abstract
Long-term assistants often receive a standing instruction once, encounter a situation in which it should not trigger, and later receive new evidence that should trigger it again. We introduce SIRRA (Standing Instruction Reapplication and Return Assessment), a benchmark for this apply–withhold–reapply behavior in long conversational histories. Each case pairs a history in which the condition remains true with a matched history in which it temporarily becomes false; both histories share the instruction, the initial and final situations, and the decision question. We score the initial action, the temporary non-action, and the later return without repeating the instruction or stating whether it applies. Across 479 cases and nine system configurations, end-to-end success across all three decisions conceals different stage bottlenecks: retrieval systems often fail before the return, while memory systems vary in whether they lose initial application, temporary restraint, or reapplication. Among cases that get the first two decisions right, return-stage outcomes vary across configurations and return forms. Explicitly stating whether the final condition holds does not uniformly improve the return and can widen the gap between the matched histories. These results show why memory evaluations should separate an instruction's continuing validity from its current applicability and report initial use, appropriate non-use, and renewed use together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.