acceptodds
Under review as a conference paper at ICLR 2027

Fisher–Rao Internal Decision Geometry for VLM Misbehavior Monitoring

Abstract

Detecting VLM misbehavior requires identifying incorrect or unsafe answers, including those generated under adversarial inputs and distribution shifts. Output confidence can remain high for such answers and does not fully describe the prediction process. Predictions with similar confidence may respond differently to small changes in decoder states, while their final distributions can conceal revisions across layers. We investigate whether these differences provide additional evidence of failure. Fisher–Rao geometry connects state variations to changes in predictive distributions and provides a shortest-path distance for comparing predictions across layers. Based on these properties, we propose Fisher–Rao Decision Geometry (FRDG), a training-free, white-box monitor. FRDG measures the strongest local predictive change under a perturbation budget relative to the state norm. It also measures accumulated changes across late-layer predictions beyond the direct distance between the first and last predictions. These measurements capture local sensitivity and intermediate revision, respectively, and their standardized summaries are combined into a risk score. We derive the local displacement interpretation and bounds on the additional layerwise change. Across four VLMs and nine benchmark subsets, the reported results show relative improvements of up to 12.42% in macro AUROC and 15.41% in macro AP over the strongest compared baseline for each model and metric. Component ablations show that the combined score improves failure ranking over either measurement alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.