When More Caution Can Make LLMs Less Safe: Risk Control for Iterative Reasoning
Abstract
Iterative LLM systems emit candidate answers over several rounds and must decide when to act, continue, or escalate; given a budget , we seek a stopping threshold controlling the wrong-action rate , which is what "safer" means here. Greater caution can increase this risk, since a higher threshold may delay action until a later round overturns an already-correct answer, so the pointwise loss monotonicity conformal risk control (CRC) requires can fail; decomposing the risk change between nested rules says exactly when caution helps. More strongly, the full risk and automation curves, the score-process distribution and every round's accuracy do not reveal whether that premise holds: two populations can agree on all of them, yet in one no trajectory violates pointwise monotonicity and in the other all do. They differ only in how correctness at one round is coupled to correctness at another, so no non-trivial bound on the violating fraction follows from any of those curves: a failure of identification, not of estimation, which an oracle would not resolve. Empirically the difficulty grows with reasoning: on MMLU-Pro, using prefixes of one run, accuracy gains saturate while correct answers broken climb from 1.9% to 8.7% of instances and the share of calibration draws with an upward risk step rises from 33% to 90%, while the aggregate curve remains nearly flat. We therefore certify stopping policies directly, keeping calibration labels isolated from score fitting, grid and budget, which gives without monotonicity and without any assumption on the score. Iterative systems need trajectory-aware certification, because aggregate risk cannot reveal whether trajectory-level monotonicity holds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.