ASPIRE: Active Supervision for Process Information-theoretic Reward Estimation
Abstract
Outcome-based Reward Models (ORMs) have advanced safety alignment for Multimodal Large Language Models (MLLMs), but their coarse feedback only indicates whether a response is safe and provides limited information about where unsafe behaviors emerge. Process Reward Models (PRMs) offer step-level supervision for safety reasoning, yet their development is constrained by the lack of fine-grained datasets, standardized evaluation protocols, and efficient training strategies. In this work, we propose a comprehensive framework for Safety PRMs. We introduce a six-class safety taxonomy, an automated step-wise annotation pipeline for generating process-level supervision, and **ProSafeBench**, a human-annotated benchmark for evaluating safety risk localization. To improve training efficiency under limited supervision, we further propose **ASPIRE** (**A**ctive **S**upervision for **P**rocess **I**nformation-theoretic **R**eward **E**stimation), which combines an information-bottleneck-inspired regularizer with uncertainty-guided example selection to identify informative training steps. Experiments show that ASPIRE achieves stronger Safety PRM performance with 25.3% fewer labeled training steps. It reaches 81.13% Global Step Accuracy on ProSafeBench and improves downstream policy optimization by reducing safety attack success rates while maintaining general multimodal capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.