acceptodds
Under review as a conference paper at ICLR 2027

SCAS: Symmetrical Concept-space Activation Steering for Joint Hallucination Mitigation and Capability Enhancement in VLMs

Abstract

A fundamental challenge in Vision-Language Models (VLMs) lies in the trade-off between mitigating hallucinations and preserving general reasoning capabilities, as existing safety interventions often degrade model utility or introduce severe biases. In this paper, we propose **SCAS** (**S**ymmetrical **C**oncept-space **A**ctivation **S**teering), a lightweight, train-free framework that achieves dual enhancement in both safety and utility. Utilizing only 100 calibration samples, SCAS extracts distinct semantic bases for factual groundedness () and hallucination (). By orthogonalizing these bases, SCAS ensures decoupled, bidirectional steering of intermediate hidden states (), simultaneously suppressing hallucination pathways while reinforcing core reasoning representations. Evaluations on Qwen3-VL-2B, Qwen3-VL-8B, and InternVL3.5-4B demonstrate that SCAS effectively addresses the safety-utility conflict: it significantly reduces visual hallucinations on RePoPe and AMBER, while consistently improving general capabilities on MME and MMBench. Furthermore, on generative captioning, SCAS avoids the decoding collapse suffered by prior methods, maintaining rich descriptive coverage while maintaining safety. This confirms SCAS as a promising solution for co-promoting VLM reliability and general intelligence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.