Continuous Oversight, Selective Deliberation: Event-Triggered Metareasoning for Judgment Revision
Abstract
Reliable judgment requires ongoing scrutiny while maintaining a responsive decision process: a strong reviewer cannot examine every request, so a system must decide which initial predictions are worth reviewing. Drawing on the metareasoning distinction between monitoring and control, we propose MONITOR, which scores every initial judgment and, under a prescribed review budget, triggers review for the requests with the highest expected correctness gain; the gain is predicted from pre-review information and learned from paired initial and reviewed outcomes. The monitor thus decides whether to investigate further, and the reviewer decides whether to change the outcome. We study two review settings: independent review of the original input, and review conditioned on the initial judgment and its rationale. We evaluate a single shared allocation rule for LLM-to-LLM and statistical-model-to-LLM review, extending the rule beyond LLM pairs, and compare both review settings on three judge benchmarks. Selection improves on random review at matched rates: at 15% review on HaluEval, MONITOR reaches 98.40% accuracy under independent review and 98.80% with the initial judgment, compared with 94.11% and 94.26% for random review. Independent review matches or exceeds full-review accuracy with few calls: 5% review on HaluEval yields 97.80% accuracy versus 97.00% for full review, while 20% on RAGTruth yields 87.50% versus 86.25%. In a single-GPU deployment on 128 HaluEval requests, selective review reaches 92.97% accuracy with 21.09% strong-model calls, compared with 92.19% for direct strong-model use; on 16 timed requests, the nominal 15% policy matches direct strong-model accuracy (87.50%) with mean final-decision times of 3.420 versus 9.175 seconds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.