acceptodds
Under review as a conference paper at ICLR 2027

MultiView Safe Decoding:Composing Cross-Horizon Evidence for Safer LLM Generation

Abstract

Lookahead can reveal several useful views of a candidate action, yet a decision based on one view can miss the implications of another. We introduce MultiView Safe Decoding (MVSD), a controller that augments an existing lookahead ranker with evidence composed across horizons. Two shared heads reuse its trace-wise features to predict absolute and pairwise executed losses, and joint calibration turns their predictions into constraints on one executed-loss vector. On a fixed panel, the controller chooses the highest-ranked candidate whose combined bounds satisfy loss and candidate-relative regret tolerances, retaining the original action when none qualifies. We establish marginal validity after action and horizon selection and give a construction in which two views jointly select an action chosen by neither view alone. A baseline-relative improvement bound connects this mechanism to changes in executed loss. Our evaluation separates benchmark performance, decision gains, evidence contributions, and computational cost, with matched-information derived heads and learned fusion as controls. On four safety benchmarks (HarmBench, SALAD-Bench, StrongREJECT, JustEval), MVSD achieves 92.92% SALAD safety and strong multi-metric performance, improving over the reference baseline by +0.48pp. However, this comes at 5 computational cost versus simpler methods. Best-of- remains superior on latency (1.46% HB ASR vs. 1.67%), while InferAligner excels on continuous harmfulness (SR: 0.01750 vs. 0.02559). We identify three failure patterns through qualitative analysis: benign edge case over-refusal (36%), inconsistent multi-view signals on creative requests (44%), and computational waste on all-safe candidate sets (20%). The framework makes cross-horizon composition explicit, demonstrates when it provides value, and honestly characterizes its limitations and trade-offs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.